Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “bioinformatic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Mining sequence annotation databanks for association patterns.

MOTIVATION: Millions of protein sequences currently being deposited to sequence databanks will never be annotated manually. Similarity-based annotation generated by automatic software pipelines unavoidably contains spurious assignments due to the imperfection of bioinformatics methods. Examples of such annotation errors include over- and underpredictions caused by the use of fixed recognition thresholds and incorrect annotations caused by transitivity based information transfer to unrelated proteins or transfer of errors already accumulated in databases. One of the most difficult and timely challenges in bioinformatics is the development of intelligent systems aimed at improving the quality of automatically generated annotation. A possible approach to this problem is to detect anomalies in annotation items based on association rule mining. RESULTS: We present the first large-scale analysis of association rules derived from two large protein annotation databases-Swiss-Prot and PEDANT-and reveal novel, previously unknown tendencies of rule strength distributions. Most of the rules are either very strong or very weak, with rules in the medium strength range being relatively infrequent. Based on dynamics of error correction in subsequent Swiss-Prot releases and on our own manual analysis we demonstrate that exceptions from strong rules are, indeed, significantly enriched in annotation errors and can be used to automatically flag them. We identify different strength dependencies of rules derived from different fields in Swiss-Prot. A compositional breakdown of association rules generated from PEDANT in terms of their constituent items indicates that most of the errors that can be corrected are related to gene functional roles. Swiss-Prot errors are usually caused by under-annotation owing to its conservative approach, whereas automatically generated PEDANT annotation suffers from over-annotation. AVAILABILITY: All data generated in this study are available for download and browsing at http://pedant.gsf.de/ARIA/index.htm.

Conserved Sequence↗

A highly conserved intraspecies homolog of the Saccharomyces cerevisiae elongation factor-3 encoded by the HEF3 gene.

A paralog (intraspecies homolog) of the Saccharomyces cerevisiae YEF3 gene, encoding elongation factor-3, has been sequenced in the course of the yeast genome project, and identified by database searching; this gene has been designated HEF3. Bioinformatic and Northern blot analysis indicate that the HEF3 gene is not expressed during vegetative growth. Deletion of the HEF3 gene reveals no growth defects, nor any defects in mating or sporulation. A high copy 2 mu clone of HEF3 was constructed, and was shown to be unable to complement a null allele of yef3. Finally, an in vitro assay for ribosome-stimulated ATPase activity was performed with isogenic HEF3 and delta hef3 strains; no difference in biochemical activity could be detected in these strains. From these results, we conclude that the HEF3 gene does not encode a functional homolog of YEF3.

Adenosine Triphosphatases↗

Website update: The UK Crop Plant Bioinformatics Network (UK CropNet).

UK CropNet currently provides a range of databases (and database-mining tools) to the plant community that are all freely accessible through our website (http://ukcrop.net/). Recent upgrades have meant that we can now expand the range of available facilities (e.g. addition of new databases) whilst also strengthening and improving access to existing services (e.g. providing a BLAST search facility against sequences in our databases). This article will briefly outline these and other new developments in our service.

Computational Biology↗

Identification of conserved regulatory elements by comparative genome analysis.

BACKGROUND: For genes that have been successfully delineated within the human genome sequence, most regulatory sequences remain to be elucidated. The annotation and interpretation process requires additional data resources and significant improvements in computational methods for the detection of regulatory regions. One approach of growing popularity is based on the preferential conservation of functional sequences over the course of evolution by selective pressure, termed 'phylogenetic footprinting'. Mutations are more likely to be disruptive if they appear in functional sites, resulting in a measurable difference in evolution rates between functional and non-functional genomic segments. RESULTS: We have devised a flexible suite of methods for the identification and visualization of conserved transcription-factor-binding sites. The system reports those putative transcription-factor-binding sites that are both situated in conserved regions and located as pairs of sites in equivalent positions in alignments between two orthologous sequences. An underlying collection of metazoan transcription-factor-binding profiles was assembled to facilitate the study. This approach results in a significant improvement in the detection of transcription-factor-binding sites because of an increased signal-to-noise ratio, as demonstrated with two sets of promoter sequences. The method is implemented as a graphical web application, ConSite, which is at the disposal of the scientific community at http://www.phylofoot.org/. CONCLUSIONS: Phylogenetic footprinting dramatically improves the predictive selectivity of bioinformatic approaches to the analysis of promoter sequences. ConSite delivers unparalleled performance using a novel database of high-quality binding models for metazoan transcription factors. With a dynamic interface, this bioinformatics tool provides broad access to promoter analysis with phylogenetic footprinting.

Algorithms↗

CINEMA--a novel colour INteractive editor for multiple alignments.

CINEMA is a new editor for manipulating and generating multiple sequence alignments. The program provides both an interface to existing databases of alignments on the Internet and a tool for constructing and modifying alignments locally. It is written in Java, so executable code will run on most major desktop platforms without modification. The implementation is highly flexible, so the applet can be easily customised with additional functions; and the object classes are reusable, promoting rapid development of program extensions. Formerly, such extended functionality might have been provided via browser plug-ins, which have to be downloaded and installed on every client before loading data. Now, for the first time, an applet is available that allows interactive client-side processing of an alignment, which can then be stored or processed automatically on the server. The program is embedded in a comprehensive help file and is accessible both as a stand-alone tool on UCL's Bioinformatics Server; http:/(/)www.biochem.ucl.ac.uk/bsm/dbbrowser+ ++/CINEMA2.02/, and as an integral part of the PRINTS protein fingerprint database. Exploitation of such novel technologies revolutionises the way users may interact with public databases in the future: bioinformatics centres need not simply provide data, but are now able to offer the means by which information is visualised and manipulated, without the requirement for users to install software.

Color Perception↗

MAIZEWALL. Database and developmental gene expression profiling of cell wall biosynthesis and assembly in maize.

An extensive search for maize (Zea mays) genes involved in cell wall biosynthesis and assembly has been performed and 735 sequences have been centralized in a database, MAIZEWALL (http://www.polebio.scsv.ups-tlse.fr/MAIZEWALL). MAIZEWALL contains a bioinformatic analysis for each entry and gene expression data that are accessible via a user-friendly interface. A maize cell wall macroarray composed of a gene-specific tag for each entry was also constructed to monitor global cell wall-related gene expression in different organs and during internode development. By using this macroarray, we identified sets of genes that exhibit organ and internode-stage preferential expression profiles. These data provide a comprehensive fingerprint of cell wall-related gene expression throughout the maize plant. Moreover, an in-depth examination of genes involved in lignin biosynthesis coupled to biochemical and cytological data from different organs and stages of internode development has also been undertaken. These results allow us to trace spatially and developmentally regulated, putative preferential routes of monolignol biosynthesis involving specific gene family members and suggest that, although all of the gene families of the currently accepted monolignol biosynthetic pathway are conserved in maize, there are subtle differences in family size and a high degree of complexity in spatial expression patterns. These differences are in keeping with the diversity of lignified cell types throughout the maize plant.

Cell Wall↗

Integration of bioInformatics tools at the National University of Singapore (NUS).

In the past decade "Big Science" such as the Genome Project has generated an enormous amount of data in the life sciences. Concurrently, the synergy of this project with existing research has quickened the pace of biological discovery. But the major drawback that is beginning to be felt worldwide is the primitive level of organisation in the data accumulated. Without a proper framework or knowledge scaffold to hang and interconnect the various bits of data and information, the national knowledge-to-data ratio is declining rapidly. We are trying to serve a solution to this enigma by providing a World Wide Web (WWW) interface to Biosoftware and at the same time have come up with a database integration tool that can query heterogeneous, geographically scattered and disparate databases simultaneously. In this report we will talk about BioInformatics in general with specific reference to BioInformatics Centre (BIC) at the National University of Singapore.

Computational Biology↗

Structural proteomics of the poxvirus family.

Recent concerns over the potential use of variola virus-commonly known as smallpox-and other orthopox viruses as weapons of bioterrorism have increased research efforts towards creating new antiviral drugs and safer more effective vaccines. Here we introduce a new resource for structural information of poxvirus proteins: the poxvirus proteomics database (PPDB). In the PPDB, we leverage recently developed bioinformatics structure prediction tools on a genomic scale and provide results in a publicly accessible format. The current version of the system contains both experimentally determined and predicted information about protein structural features, such as secondary structure and relative solvent accessibility, as well as tertiary structure and homology information. The system is automated to read the primary sequences from the database, produce the new information for each sequence, and update the database monthly and as new tools are incorporated. The PPDB contains detailed information on the open reading frames (ORFs) in the Copenhagen strain of the vaccinia virus genome. The contents of the PPDB can be accessed through a simple web interface. Inclusion of additional poxvirus genomes in the PPDB is in progress. The PPDB has an upward scalable informatics infrastructure that can readily be applied to viral, bacterial, as well as eukaryotic genomes.

Antiviral Agents↗

Persistent gaps and errors in reference databases impede ecologically meaningful taxonomy assignments in 18S rRNA studies: a case study of terrestrial and marine nematodes.

In metabarcoding studies, Linnaean taxonomy assignments of Operational Taxonomic Units (OTUs) or Amplicon Sequence Variants (ASVs) underpin many downstream bioinformatics analyses and ecological interpretations of environmental DNA (eDNA) datasets. However, public molecular databases (i.e., SILVA, EUKARYOME, BOLD) for most microbial metazoan phyla (nematodes, tardigrades, kinorhynchs, etc.) are sparsely populated, negatively impacting our ability to assign ecologically meaningful taxonomy to these understudied groups. Additionally, the choice of bioinformatics parameters and computational algorithms can further impact the accuracy of eDNA taxonomy assignments. Here, we use two in-silico datasets to show that taxonomy assignments using the 18S rRNA gene can be dramatically improved by curating Linnaean taxonomy strings associated with each reference sequence and closing phylogenetic gaps by improving taxon sampling. Using free-living nematodes as a case study, we applied two commonly used taxonomy assignment algorithms (BLAST+ and the QIIME2 Naïve Bayes classifier) across six iterations of the SILVA 138 reference database to evaluate the precision and accuracy of taxonomy assignments. The BLAST+ top hit with a 90% sequence similarity cutoff often returned the highest percentage of correctly assigned taxonomy at the genus level, and the QIIME2 Naïve Bayes classifier performed similarly well when paired with a reference database containing corrected taxonomy strings. Our results highlight the urgent need for phylogenetically-informed expansions of public reference databases (encompassing both genomes and common gene markers), focused on poorly sampled lineages which are now robustly recovered via eDNA metabarcoding approaches. Additional taxonomy curation efforts should be applied to popular reference databases such as SILVA, and taxon sampling could be rapidly improved by more frequent incorporation of newly published GenBank sequences linked to genus and/or species level identifications.

18S rRNA metabarcoding↗

Identification of Heterodera glycines (soybean cyst nematode [SCN]) cDNA sequences with high identity to those of Caenorhabditis elegans having lethal mutant or RNAi phenotypes.

The soybean cyst nematode (SCN; Heterodera glycines) is a devastating obligate parasite of Glycine max (soybean) causing one billion dollars in losses to the US economy per year and over ten billion dollars in losses worldwide. While much is understood about the pathology of H. glycines, its genome sequence is not well characterized or fully sequenced. We sought to create bioinformatic tools to mine the H. glycines nucleotide database. One way is to use a comparative genomics approach by anchoring our analysis with an organism, like the free-living nematode Caenorhabditis elegans. Unlike H. glycines, the C. elegans genome is fully sequenced and is well characterized with a number of lethal genes identified through experimental methods. We compared an EST database of H. glycines with the C. elegans genome. Our goal was identifying genes that may be essential for H. glycines survival and would serve as an automated pipeline for RNAi studies to both study and control H. glycines. Our analysis yielded a total of nearly 8334 conserved genes between H. glycines and C. elegans. Of these, 1508 have lethal phenotypes/phenocopies in C. elegans. RNAi of a conserved ribosomal gene from H. glycines (Hg-rps-23) yielded dead and dying worms as shown by positive Sytox fluorescence. Endogenous Hg-rps-23 exhibited typical RNA silencing as shown by RT-PCR. However, an unrelated gene Hg-unc-87 did not exhibit RNA silencing in the Hg-rps-23 dsRNA-treated worms, demonstrating the specificity of the silencing.

Animals↗

Metabolomics in practice: emerging knowledge to guide future dietetic advice toward individualized health.

The profession of dietetics can take an increasingly prominent role in managing health and patient care as clinicians gain access to three new resources: detailed information about the metabolic status of healthy individual clients, metabolic knowledge about the relationships between metabolite abundances and health, and bioinformatics tools that link clients' metabolism to their present and future health status. The current use of single biomarkers as indicators of disease will be replaced by comprehensive profiling of individual metabolites linked to an understanding of health and human metabolism--the emerging science now known as metabolomics. Industrial and academic initiatives are currently developing the analytical and bioinformatic technologies needed to assemble the quantitative reference databases of metabolites as the metabolic analog of the human genome. With these in place, dietetics professionals will be able to assess both the current health status of individuals and predict their health trajectories. Another important role for dietetics professionals will be to assist in the development of the tools and their application in predicting how an individual's specific metabolic pattern can be changed by diet, drugs, and lifestyle, with the goal of improving health and preventing the development of chronic diseases.

Biomarkers↗

In need of high-throughput behavioral systems.

One of the current major bottlenecks in drug discovery is in vivo testing of candidate drugs in behavioral paradigms in normal or genetically altered mice. This testing is essential in discovering gene function and predicting potential efficacy of CNS drugs in humans. New efforts in the biotech community aim to alleviate this bottleneck by developing higher-throughput systems of behavioral, neurological and physiological analyses. Together with large pharmacological databases, equipped with state-of-the-art bioinformatic and/or data-mining algorithms, these systems will provide rapid and accurate indices of the therapeutic potential of novel drugs. By providing a substantial increase in the speed of behavioral testing, new high-throughput systems will facilitate current behavioral research with faster, more reliable approaches. Furthermore, screening whole drug-libraries and comparing the profiles of novel compounds to those of known compounds will facilitate the discovery of novel drugs. Target validation will also become more efficient with the fast characterization of novel mutant mice.

Animals↗

Bioinformatics, functional genomics, and proteomics study of Bacillus sp.

The ability of bioinformatics to characterize genomic and proteomic sequences from bacteria Bacillus sp. for prediction of genes and proteins has been evaluated. Genomics coupling with proteomics, which is relied on integration of the significant advances recently achieved in two-dimensional (2-D) electrophoretic separation of proteins and mass spectrometry (MS), are now important and high throughput techniques for qualifying and analyzing gene and protein expression, discovering new gene or protein products, and understanding of gene and protein functions including post-genomic study. In addition, the bioinformatics of Bacillus sp. is embraced into many databases that will facilitate to rapidly search the information of Bacillus sp. in both genomics and proteomics. It is also possible to highlight sites for post-translational modifications based on the specific protein sequence motifs that play important roles in the structure, activity and compartmentalization of proteins. Moreover, the secreted proteins from Bacillus sp. are interesting and widely used in many applications especially biomedical applications that are the highly advantages for their potential therapeutic values.

Bacillus↗

The involvement of lncRNA EMSLR in the disulfidptosis and progression of endometrial carcinoma.

The incidence of endometrial cancer (EC) continues to rise. Disulfidptosis, a novel form of cell death, may represent a potential therapeutic target in EC. Through bioinformatic analysis of The Cancer Genome Atlas (TCGA) database, E2F1 mRNA-stabilizing lncRNA (EMSLR) was identified as a lncRNA related to disulfidptosis in EC. Functional assays, including cell proliferation and xenograft assays, demonstrated that knockdown of EMSLR significantly impeded EC cell proliferation, whereas overexpression of EMSLR promoted cell viability. Additionally, EMSLR was found to be associated with glucose uptake and NADPH production in glucose-restricted culture conditions. Moreover, downregulation of EMSLR markedly increased cell death and induced cytoskeletal collapse under glucose deprivation, as evidenced by F-actin and cell death staining. Notably, we observed a strong correlation between EMSLR and the c-MYC-GLUT1 pathway. Mechanistically, EMSLR was found to mediate the expression and nuclear translocation of c-MYC, thereby regulating the progression of EC and its associated disulfidptosis. In conclusion, EMSLR is identified as a disulfidptosis-related gene in endometrial cancer. Elucidating the function and molecular mechanisms of EMSLR in EC presents a promising avenue for therapeutic intervention in patients.

Female↗

Prospective use of DNA microarrays for evaluating renal function and disease.

At the forefront of the revolution in human genomics is DNA microarray technology, which evaluates expression levels or genotypes of thousands of genes simultaneously, by means of miniaturization and parallel processing. Furthermore, advances in bioinformatics will result in the creation of large databases, which will require complex software programming for structural analysis. Over the next decade, DNA microarrays, combined with sophisticated informatics and genomic databases, will provide molecular fingerprints of disease processes and prognoses. This review provides an update on DNA microarray technology and its application to renal diseases.

DNA↗

Genomic tools and cDNA derived markers for butterflies.

The Lepidoptera have long been used as examples in the study of evolution, but some questions remain difficult to resolve due to a lack of molecular genetic data. However, as technology improves, genomic tools are becoming increasingly available to tackle unanswered evolutionary questions. Here we have used expressed sequence tags (ESTs) to develop genetic markers for two Müllerian mimic species, Heliconius melpomene and Heliconius erato. In total 1363 ESTs were generated, representing 330 gene objects in H. melpomene and 431 in H. erato. User-friendly bioinformatic tools were used to construct a nonredundant database of these putative genes (available at http://www.heliconius.org), and annotate them with blast similarity searches, InterPro matches and Gene Ontology terms. This database will be continually updated with EST sequences for the Papilionideae as they become publicly available, providing a tool for gene finding in the butterflies. Alignments of the Heliconius sequences with putative homologues derived from Bombyx mori or other public data sets were used to identify conserved PCR priming sites, and develop 55 markers that can be amplified from genomic DNA in both H. erato and H. melpomene. These markers will be used for comparative linkage mapping in Heliconius and will have applications in other phylogenetic and genomic studies in the Lepidoptera.

Adaptation, Biological↗

Expanding vaginal microbiome pangenomes via a custom MIDAS database reveals Lactobacillus crispatus accessory genes associated with cervical dysplasia.

The vaginal microbiome plays a central role in reproductive health. Vaginal microbiome dysbiosis is associated with many adverse reproductive health outcomes, but most studies have focused on associations at the species level. The potential contribution of intraspecies microbial variation, especially gene content differences across bacterial strains, remains underexplored in reproductive health contexts. The Metagenomic Intra-Species Diversity Analysis (MIDAS) framework enables such analyses, but depends on comprehensive reference databases. We constructed a MIDAS-compatible pangenome database from over 18,000 genomes in the Vaginal Microbiome Genome Collection (VMGC). Compared to the Genome Taxonomy Database (GTDB)-derived reference, the VMGC-derived database expanded the pangenomes of prevalent vaginal species, better capturing vaginal-specific intraspecies diversity. Applying this database to vaginal samples from a cervical dysplasia cohort, we identified 13 Lactobacillus crispatus accessory genes significantly associated with cervical dysplasia, including a HicAB toxin-antitoxin system, three transcriptional regulators, and three phage-derived genes. These findings highlight the utility of body site-specific reference resources and shotgun metagenomic sequencing for uncovering intraspecies microbial variation relevant to reproductive health.IMPORTANCEThe vaginal microbiome plays a critical role in reproductive health, and different bacteria from the same species can carry different genes that influence how the strains interact with the host and other microbes. These strain-level differences are often overlooked when microbiomes are analyzed only at the species level. Existing genomic reference databases are heavily biased toward gut and environmental bacteria, leaving the genetic diversity of vaginal microbes understudied. We built a specialized reference database from over 18,000 vaginal bacterial genomes that better reflects this diversity. We then applied this resource to quantify gene-level variation in vaginal samples from a cervical dysplasia cohort. Focusing on Lactobacillus crispatus, a prevalent and often beneficial vaginal species, we identified 13 genes that were more common in women with cervical dysplasia than in controls. This work demonstrates that body site-specific genomic resources are essential for uncovering strain-level bacterial differences relevant to reproductive health.

Lactobacillus crispatus↗

RiboSubstrates: a web application addressing the cleavage specificities of ribozymes in designated genomes.

BACKGROUND: RNA-dependent gene silencing is becoming a routine tool used in laboratories worldwide. One of the important remaining hurdles in the selection of the target sequence, if not the most important one, is the designing of tools that have minimal off-target effects (i.e. cleaves only the desired sequence). Increasingly, in the current dawn of the post-genomic era, there is a heavy reliance on tools that are suitable for high-throughput functional genomics, consequently more and more bioinformatic software is becoming available. However, to date none have been designed to satisfy the ever-increasing need for the accurate selection of targets for a specific silencing reagent. RESULTS: In order to overcome this hurdle we have developed RiboSubstrates http://www.riboclub.org/ribosubstrates. This integrated bioinformatic software permits the searching of a cDNA database for all potential substrates for a given ribozyme. This includes the mRNAs that perfectly match the specific requirements of a given ribozyme, as well those including Wobble base pairs and mismatches. The results generated allow rapid selection of sequences suitable as targets for RNA degradation. The current web-based RiboSubstrates version permits the identification of potential gene targets for both SOFA-HDV ribozymes and for hammerhead ribozymes. Moreover, a minimal template for the search of siRNAs is also available. This flexible and reliable tool is easily adaptable for use with any RNA tool (i.e. other ribozymes, deoxyribozymes and antisense), and may use the information present in any cDNA bank. CONCLUSION: RiboSubstrates should become an essential step for all, even including "non-RNA biologists", who endeavor to develop a gene-inactivation system.

Animals↗