Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

Synchrotron and neutron techniques in biological crystallography.

Synchrotron radiation (SR) techniques are continuously pushing the frontiers of wavelength range usage, smaller crystal sample size, larger protein molecular weight and complexity, as well as better diffraction resolution. The new research specialism of probing functional states directly in crystals, via time-resolved Laue and freeze trapping structural studies, has been developed, with a range of examples, based on research stretching over some 20 years. Overall, SR X-ray biological crystallography is complemented by neutron protein crystallographic studies aimed at cases where much more complete hydrogen details are needed involving synergistic developments between SR and neutron Laue methods. A big new potential exists in harnessing genome databases for targeting of new proteins for structural study. Structural examples in this tutorial review illustrate new chemistry learnt from biological macromolecules.

Crystallography, X-Ray↗

Informatics--genome and genetic databases.

Databases for biologists are becoming increasingly important. Some of these can be regarded as 'core' resources, such as the bibliographic databases, whereas others are of greater interest to specialists. As comparative genomics develops, however, even databases limited in their scope (e.g. to a single organism) are of great interest to a wider community.

Animals↗

Automated characterization of potentially active retroid agents in the human genome.

Retroid agents are genomes that encode the reverse transcriptase (RT) and replicate by way of an RNA intermediate. Some retroid agents are implicated in disease via insertional mutagenesis, while others have been found to encode proteins essential to primate reproduction or provide regulatory sequences for host cell processes. The Genome Parsing Suite (GPS), a generic multistep automated process, was developed to characterize all RT-like sequences in the human genome database and to annotate the gene complement of the retroid agents that encode these sequences. In this report the GPS analyzes all significant WU-tBLASTn hits returned for 30 representative RT queries. A total of 128,779 unique RT signals were identified, and 7594 of these were retrieved by RTs not previously reported in the human genome. We have identified 9652 full-length long interspersed nuclear elements (LINEs). Only 159 LINEs are without stop codons or frameshifts.

Amino Acid Motifs↗

Expression of a novel neuropeptide, NVGTLARDFQLPIPNamide, in the larval and adult brain of Drosophila melanogaster.

Advances in mass spectrometry and the availability of genomic databases made it possible to determine the peptidome or peptide content of a specific tissue. Peptidomics by nanoflow capillary liquid chromatography tandem mass spectrometry of an extract of 50 larval Drosophila brains, yielded 28 neuropeptides. Eight were entirely novel and encoded by five not yet annotated genes; only two genes had a homologue in the Anopheles gambiae genome. Seven of the eight peptides did not show relevant sequence homology to any known peptide. Therefore, no evidence towards the physiological role of these 'orphan' peptides was available. We identified one of the eight peptides, IPNamide, in an extract of the Drosophila adult brain as well. Next, specific antisera were raised to reveal the distribution pattern of IPNamide and other peptides from the same precursor, in larval and adult brains by means of whole-mount immunocytochemistry and confocal microscopy. IPNamide immunoreactivity is abundantly present in both stages and a striking similarity was found between the distribution patterns of IPNamide and TPAEDFMRFamide, a member of the FMRFamide peptide family. Based on this distribution pattern, IPNamide might be involved in phototransduction, in processing sensory stimuli, as well as in controlling the activity of the oesophagus.

Amino Acid Sequence↗

A comparative hybridization analysis of yeast DNA with Paramecium parafusin- and different phosphoglucomutase-specific probes.

Molecular probes designed for the parafusin (PFUS), the Paramecium exocytic-sensitive phosphoglycoprotein, gave distinct hybridization patterns in Saccharomyces cerevisiae genomic DNA when compared with different phosphoglucomutase specific probes. These include two probes identical to segments of yeast phosphoglucomutase (PGM) genes 1 and 2. Neither of the PGM probes revealed the 7.4 and 5.9 kb fragments in Bgl II-cut yeast DNA digest detected with the 1.6 kb cloned PFUS cDNA and oligonucleotide constructed to the PFUS region (insertion 3--I-3) not found in other species. PCR amplification with PFUS-specific primers generated yeast DNA-species of the predicted molecular size which hybridized to the I-3 probe. A search of the yeast genome database produced an unassigned nucleotide sequence that showed 55% identity to parafusin gene and 37% identity to PGM2 (the major isoform of yeast phosphoglucomutase) within the amplified region.

Animals↗

Phylogenomic analysis of non-ribosomal peptide synthetases in the genus Aspergillus.

Fungi from the genus Aspergillus are important saprophytes and opportunistic human fungal pathogens that contribute in these and other diverse ways to human well-being. Part of their impact on human well-being stems from the production of small molecular weight secondary metabolites, which may contribute to the ability of these fungi to cause invasive fungal infections and allergic diseases. In this study, we identified one group of enzymes responsible for secondary metabolite production in five Aspergillus species, the non-ribosomal peptide synthetases (NRPS). Hidden Markov models were used to search the genome databases of A. fumigatus, A. flavus, A. terreus, A. nidulans, and A. oryzae for domains conserved in NRPS proteins. A genealogy of adenylation domains was utilized to identify orthologous and unique NRPS among the Aspergillus species examined, as well as gain an understanding of the potential evolution of Aspergillus NRPS. mRNA abundance of the 14 NRPS identified in the A. fumigatus genome was analyzed using real-time reverse transcriptase PCR in different environmental conditions to gain a preliminary understanding of the possible functions of the NRPSs' peptide products. Our results suggest that Aspergillus species contain conserved and unique NRPS genes with a complex evolutionary history. This result suggests that the genus Aspergillus produces a substantial diversity of non-ribosomally synthesized peptides. Further analysis of these genes and their peptide products may identify important roles for secondary metabolites produced by NRPS in Aspergillus physiology, ecology, and fungal pathogenicity.

Amino Acid Sequence↗

Functional genomics of hepatocellular carcinoma.

The majority of DNA-microarray based gene expression profiling studies on human hepatocellular carcinoma (HCC) has focused on identifying genes associated with clinicopathological features of HCC patients. Although notable success has been achieved, this approach still faces significant challenges due to the heterogeneous nature of HCC (and other cancers) as well as the many confounding factors embedded in gene expression profile data. However, these limitations are being overcome by improved bioinformatics and sophisticated analyses. Also, application of cross comparison of multiple gene expression data sets from human tumors and animal models are facilitating the identification of critical regulatory modules in the expression profiles. The success of this new experimental approach, comparative functional genomics, suggests that integration of independent data sets will enhance our ability to identify key regulatory elements in tumor development. Furthermore, integrating gene expression profiles with data from DNA sequence information in promoters, array-based CGH, and expression of non-coding genes (i.e., microRNAs) will further increase the reliability and significance of the biological and clinical inferences drawn from the data. The pace of current progress in the cancer profiling field, combined with the advances in high-throughput technologies in genomics and proteomics, as well as in bioinformatics, promises to yield unprecedented biological insights from the integrative (or systems) analysis of the combined cancer genomics database. The predicted beneficial impact of this "new integrative biology" on diagnosis, treatment and prevention of liver cancer and indeed cancer in general is enormous.

Animals↗

The maize INDETERMINATE1 flowering time regulator defines a highly conserved zinc finger protein family in higher plants.

BACKGROUND: The maize INDETERMINATE1 gene, ID1, is a key regulator of the transition to flowering and the founding member of a transcription factor gene family that encodes a protein with a distinct arrangement of zinc finger motifs. The zinc fingers and surrounding sequence make up the signature ID domain (IDD), which appears to be found in all higher plant genomes. The presence of zinc finger domains and previous biochemical studies showing that ID1 binds to DNA suggests that members of this gene family are involved in transcriptional regulation. RESULTS: Comparison of IDD genes identified in Arabidopsis and rice genomes, and all IDD genes discovered in maize EST and genomic databases, suggest that ID1 is a unique member of this gene family. High levels of sequence similarity amongst all IDD genes from maize, rice and Arabidopsis suggest that they are derived from a common ancestor. Several unique features of ID1 suggest that it is a divergent member of the maize IDD family. Although no clear ID1 ortholog was identified in the Arabidopsis genome, highly similar genes that encode proteins with identity extending beyond the ID domain were isolated from rice and sorghum. Phylogenetic comparisons show that these putative orthologs, along with maize ID1, form a group separate from other IDD genes. In contrast to ID1 mRNA, which is detected exclusively in immature leaves, several maize IDD genes showed a broad range of expression in various tissues. Further, Western analysis with an antibody that cross-reacts with ID1 protein and potential orthologs from rice and sorghum shows that all three proteins are detected in immature leaves only. CONCLUSION: Comparative genomic analysis shows that the IDD zinc finger family is highly conserved among both monocots and dicots. The leaf-specific ID1 expression pattern distinguishes it from other maize IDD genes examined. A similar leaf-specific localization pattern was observed for the putative ID1 protein orthologs from rice and sorghum. These similarities between ID1 and closely related genes in other grasses point to possible similarities in function.

Amino Acid Sequence↗

Tandem genes of Chlamydia psittaci that encode proteins localized to the inclusion membrane.

Chlamydiae are obligate intracellular bacteria that replicate within a non-acidified vacuole, termed an inclusion. To identify chlamydial proteins that are unique to the intracellular phase of the life cycle, a lambda expression library of Chlamydia psittaci DNA was differentially screened with convalescent antisera from infected guinea pigs and antisera directed at formalin-fixed purified chlamydial elementary bodies (EBs). One library clone was identified that harboured two open reading frames (ORFs) with coding potential for similar-sized proteins of approximately 20 kDa. These proteins were subsequently termed IncB and IncC. Sequencing of the cloned insert revealed a strong Escherichia coli-like promoter sequence immediately upstream of incB and a 36nt intergenic region between the ORFs. Sequence analysis of the region upstream of incB and incC revealed two ORFs that had strong homologies to an amino acid transporter and a sodium-dependent transporter. Immunoblotting with antisera directed at IncB or IncC demonstrated that these proteins are present in C. psittaci-infected HeLa cells but are absent or below the level of detection in purified EBs. Reverse transcriptase-polymerase chain reactions provided evidence that incB and incC are transcribed in an operon. Immunofluorescence microscopy demonstrated that IncB and IncC are each localized to the inclusion membrane of infected cells. No primary sequence similarity is evident between IncA, IncB or IncC, but each contains a large hydrophobic domain of similar size and character as in IncA. Analysis of the recently completed C. trachomatis serovar D genome database has revealed C. trachomatis ORFs encoding homologues to incB and incC, indicating that these genes are conserved among the chlamydiae.

Animals↗

In silico search for functionally similar proteins involved in meiosis and recombination in evolutionarily distant organisms.

Evolutionarily distant organisms have not only orthologs, but also nonhomologous proteins that build functionally similar subcellular structures. For instance, this is true with protein components of the synaptonemal complex (SC), a universal ultrastructure that ensures the successful pairing and recombination of homologous chromosomes during meiosis. We aimed at developing a method to search databases for genes that code for such nonhomologous but functionally analogous proteins. Advantage was taken of the ultrastructural parameters of SC and the conformation of SC proteins responsible for these. Proteins involved in SC central space are known to be similar in secondary structure. Using published data, we found a highly significant correlation between the width of the SC central space and the length of rod-shaped central domain of mammalian and yeast intermediate proteins forming transversal filaments in the SC central space. Basing on this, we suggested a method for searching genome databases of distant organisms for genes whose virtual proteins meet the above correlation requirement. Our recent finding of the Drosophila melanogaster CG17604 gene coding for synaptonemal complex transversal filament protein received experimental support from another lab. With the same strategy, we showed that the Arabidopsis thaliana and Caenorhabditis elegans genomes contain unique genes coding for such proteins.

Animals↗

Genomic mapping of chemokine and transforming growth factor genes in swine.

Five chemokine genes, transforming growth factors alpha, beta 2 and 3 (TGFBA, TGFB-2, and TGFB-3), interleukin 8 (IL-8), and monocyte chemoattractant protein 2 (MCP-2), were mapped to porcine linkage groups on Chromosomes 3q, 10p, 7q, 8, and 12q, respectively. Restriction fragment length polymorphisms (RFLPs) for these genes were developed by Southern blot hybridization after digestion of porcine genomic DNA with BamHI and MspI (TGFBA), BamHI and PvuII (TGFB-2), HindIII (TGFB-3), BglII (IL-8), and PstI (MCP-2) and used to genotype the USDA-MARC Swine Reference Population pigs. Sufficient informative meioses, 61 (TGFBA), 58 (TGFB-2), 28 (TGFB-3), 38 (IL-8), and 156 (MCP-2), were available to pursue two-point pairwise linkage analysis with over 1,000 existing loci in the USDA-MARC genome database to establish initial linkage (LOD > 3). Multi-point analysis with CRIMAP determined the most likely order for each new marker. The assignment of the five chemokine genes in swine concurs with previous porcine/human chromosomal homologies based on results from ZOO-FISH and chromosomal painting experiments. These findings add five new informative Type I markers within a single gene family to the swine genome and may help us understand the genetic basis for disease resistance in livestock.

Animals↗

Data mining for protein-protein interactions in invertebrate model organisms.

Well-annotated genome databases are available for many invertebrate species, notably the fruitfly, Drosophila melanogaster, and the nematode, Caenorhabditis elegans. An adequate interpretation of this information at the biological level requires the exploration of the interactions between the gene products. Knowledge of protein interactions and the components of cell signalling pathways in the fly and worm are particularly valuable as hypotheses can be rapidly tested using the powerful genetic toolkits available. Invertebrates offer additional experimental advantages when attempting to characterise protein-protein interactions (PPIs). Their relatively small genome size compared to mammals helps to reduce missed interactions due to redundancy, and their function can be addressed using forward (mutants) and reverse (RNA interference) genetics. However, the researcher looking for evidence of PPIs for a protein of interest is faced with the challenge of extracting interaction data from sources that are highly varied, such as the results of microarray experiments in the unstructured text of research papers. This challenge is greatly reduced by a range of public databases of curated information, as well as publicly available, enhanced search engines, which can provide either direct experimental evidence for a PPI, or valuable clues for generating new hypotheses.

Animals↗

Identification of a novel gene, URE2, that functionally complements a urease-negative clinical strain of Cryptococcus neoformans.

A urease-negative serotype A strain of Cryptococcus neoformans (B-4587) was isolated from the cerebrospinal fluid of an immunocompetent patient with a central nervous system infection. The URE1 gene encoding urease failed to complement the mutant phenotype. Urease-positive clones of B-4587 obtained by complementing with a genomic library of strain H99 harboured an episomal plasmid containing DNA inserts with homology to the sudA gene of Aspergillus nidulans. The gene harboured by these plasmids was named URE2 since it enabled the transformants to grow on media containing urea as the sole nitrogen source while the transformants with an empty vector failed to grow. Transformation of strain B-4587 with a plasmid construct containing a truncated version of the URE2 gene failed to complement the urease-negative phenotype. Disruption of the native URE2 gene in a wild-type serotype A strain H99 and a serotype D strain LP1 of C. neoformans resulted in the inability of the strains to grow on media containing urea as the sole nitrogen source, suggesting that the URE2 gene product is involved in the utilization of urea by the organism. Virulence in mice of the urease-negative isolate B-4587, the urease-positive transformants containing the wild-type copy of the URE2 gene, and the urease-negative vector-only transformants was comparable to that of the H99 strain of C. neoformans regardless of the infection route. Virulence of the URE2 disruption stain of H99 was slightly reduced compared to the wild-type strain in the intravenous model but was significantly attenuated in the inhalation model. These results indicate that the importance of urease activity in pathogenicity varies depending on the strains of C. neoformans used and/or the route of infection. Furthermore, this study shows that complementation cloning can serve as a useful tool to functionally identify genes such as URE2 that have otherwise been annotated as hypothetical proteins in genomic databases.

Animals↗

Identification and characterization of Thermoplasma acidophilum glyceraldehyde dehydrogenase: a new class of NADP+-specific aldehyde dehydrogenase.

Thermoacidophilic archaea such as Thermoplasma acidophilum and Sulfolobus solfataricus are known to metabolize D-glucose via the nED (non-phosphorylated Entner-Doudoroff) pathway. In the present study, we identified and characterized a glyceraldehyde dehydrogenase involved in the downstream portion of the nED pathway. This glyceraldehyde dehydrogenase was purified from T. acidophilum cell extracts by sequential chromatography on DEAE-Sepharose, Q-Sepharose, Phenyl-Sepharose and Affi-Gel Blue columns. SDS/PAGE of the purified enzyme showed a molecular mass of approx. 53 kDa, whereas the molecular mass of the native protein was 215 kDa, indicating that glyceraldehyde dehydrogenase is a tetrameric protein. By MALDI-TOF-MS (matrix-assisted laser-desorption ionization-time-of-flight MS) peptide fingerprinting of the purified protein, it was found that the gene product of Ta0809 in the T. acidophilum genome database corresponds to the purified glyceraldehyde dehydrogenase. The native enzyme showed the highest activity towards glyceraldehyde, but no activity towards aliphatic or aromatic aldehydes, and no activity when NAD+ was substituted for NADP+. Analysis of the amino acid sequence and enzyme inhibition studies indicated that this glyceraldehyde dehydrogenase belongs to the ALDH (aldehyde dehydrogenase) superfamily. BLAST searches showed that homologues of the Ta0809 protein are not present in the Sulfolobus genome. Possible differences between T. acidophilum (Euryarchaeota) and S. solfataricus (Crenarchaeaota) in terms of the glycolytic pathway are thus expected.

Aldehyde Dehydrogenase↗

Identification and expression of the Saccharomyces cerevisiae cytoplasmic tryptophanyl-tRNA synthetase gene.

The enzymes that aminoacylate tRNAs have been studied extensively and can be organized into two distinct classes based on signature sequences and the position of aminoacylation. The class I enzymes have canonical HIGH and KMSKS sequences as part of a Rossman fold nucleotide-binding site. The tryptophan-specific enzymes have been placed in class I based on analysis of the cognate genes from Escherichia coli, B. stearothermophilus, B. taurus, and Homo sapiens. An unidentified open reading frame (ORF) on Saccharomyces cerevisiae chromosome XV, HRE342, has 46% identity with the bovine tryptophanyl-tRNA synthetase and possesses the appropriate signature sequences. The predicted molecular weight of the putative HRE342 protein also closely matched the expected monomer size of the S. cerevisiae enzyme. The HRE342 ORF plus about 250 bp of 5' and 3' flanking sequence was amplified by polymerase chain reaction, cloned into a 2 mu based vector, and transformed into a host strain, S. cerevisiae JG369.3B. Nucleotide sequence analysis of the clone confirmed the presence of HRE342. Extracts from transformed yeast have a 30- to 100-fold increase in specific activity of the tryptophanyl-tRNA synthetase. An HRE342 locus in a diploid strain, PTY33XPTY44, was disrupted with a LEU2 insert. Sporulation and tetrad analysis of the HRE342::LEU2 strain demonstrated that HRE342 is an essential gene. We conclude that HRE342 is the S. cerevisiae gene encoding the cytoplasmic tryptophanyl-tRNA synthetase, WRS1. A search of the Saccharomyces Genome Database using amino acid sequences from other eukaryotic aminoacyl-tRNA synthetase suggests there is sufficient similarity to identify both class I and class II genes.

Amino Acid Sequence↗

Molecular and culture-based analyses of aerobic carbon monoxide oxidizer diversity.

Isolates belonging to six genera not previously known to oxidize CO were obtained from enrichments with aquatic and terrestrial plants. DNA from these and other isolates was used in PCR assays of the gene for the large subunit of carbon monoxide dehydrogenase (coxL). CoxL and putative coxL fragments were amplified from known CO oxidizers (e.g., Oligotropha carboxidovorans and Bradyrhizobium japonicum), from novel CO-oxidizing isolates (e.g., Aminobacter sp. strain COX, Burkholderia sp. strain LUP, Mesorhizobium sp. strain NMB1, Stappia strains M4 and M8, Stenotrophomonas sp. strain LUP, and Xanthobacter sp. strain COX), and from several well-known isolates for which the capacity to oxidize CO is reported here for the first time (e.g., Burkholderia fungorum LB400, Mesorhizobium loti, Stappia stellulata, and Stappia aggregata). PCR products from several taxa, e.g., O. carboxidovorans, B. japonicum, and B. fungorum, yielded sequences with a high degree (>99.6%) of identity to those in GenBank or genome databases. Aligned sequences formed two phylogenetically distinct groups. Group OMP contained sequences from previously known CO oxidizers, including O. carboxidovorans and Pseudomonas thermocarboxydovorans, plus a number of closely related sequences. Group BMS was dominated by putative coxL sequences from genera in the Rhizobiaceae and other alpha-PROTEOBACTERIA: PCR analyses revealed that many CO oxidizers contained two coxL sequences, one from each group. CO oxidation by M. loti, for which whole-genome sequencing has revealed a single BMS-group putative coxL gene, strongly supports the notion that BMS sequences represent functional CO dehydrogenase proteins that are related to but distinct from previously characterized aerobic CO dehydrogenases.

Aerobiosis↗

Integrating computationally assembled mouse transcript sequences with the Mouse Genome Informatics (MGI) database.

Databases of experimentally generated and computationally derived transcript sequences are valuable resources for genome analysis and annotation. The utility of such databases is enhanced when the sequences they contain are integrated with such biological information as genomic location, gene function, gene expression and phenotypic variation. We present the analysis and results of a semi-automated process of connecting transcript assemblies with highly curated biological information for mouse genes that is available through the Mouse Genome Informatics (MGI) database.

Animals↗

An open-source clinical bioinformatics pipeline for real-world NGS implementation: translating genomic variants into actionable treatment strategies in oncology.

BACKGROUND: Next-Generation Sequencing (NGS) has become a cornerstone technology in clinical practice, yet its adoption presents significant challenges. Physicians and oncologists must manage vast amounts of genome-scale data and transform it into actionable insights for complex decision-making. While commercial systems exist to synthesize data from NGS experiments into clinical reports, many are hindered by limitations such as closed-source designs that restrict transparency and customization. Additionally, some fail to leverage publicly available genomic databases, missing opportunities to integrate valuable external data. Furthermore, the rigidity of many tools in accommodating diverse NGS panels limits their applicability across varied clinical scenarios. METHODS: To address these limitations, we developed OncoReport, an open-source tool that generates comprehensive reports from NGS analyses. By integrating publicly accessible databases, OncoReport provides a robust, user-friendly environment equipped with essential tools for NGS analysis. This design aims to enhance data interpretation and support informed clinical decision-making. RESULTS: Rigorous testing has demonstrated OncoReport’s effectiveness in producing detailed, actionable reports that are clear and easy to use. By automating key aspects of the workflow, the tool significantly reduces manual effort and expedites the synthesis and interpretation of NGS results, making genomic insights more accessible to clinicians. CONCLUSION: OncoReport offers a transparent, flexible, and efficient framework for clinicians to analyze and apply genomic data in patient care. By streamlining workflows and leveraging open-source principles, it empowers healthcare professionals to make informed, data-driven decisions. OncoReport is freely available at https://oncoreport.atlas.dmi.unict.it, with source code and issue tracking on GitHub: https://github.com/knowmics-lab/oncoreport .

Humans↗