Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

IRIS: a database surveying known human immune system genes.

We have compiled an online database of known human defense genes: the Immunogenetic Related Information Source (IRIS). As of October 1, 2004, there are 1562 immune genes recorded in IRIS, representing 7% of the human genome. This resource contains searchable information including chromosomal location, sequence data, and a curated functional annotation for each entry. We used IRIS as a basis for analyzing the composition and characteristics of the immune genome, such as gene clustering, polymorphism, and relationship to disease. High protein sequence similarity correlated inversely with distance between immune genes, consistent with clustering of duplicated loci. We also found that, even though some immune genes exhibit high levels of polymorphism, such as MHC class I, the range of levels of polymorphism in immune genes is similar to that of nonimmune genes. Approximately 20% of immune genes have a known disease association. IRIS is available online at .

Databases, Genetic↗

The role of protein structure in genomics.

The genome projects produce an enormous amount of sequence data that needs to be annotated in terms of molecular structure and biological function. These tasks have triggered additional initiatives like structural genomics. The intention is to determine as many protein structures as possible, in the most efficient way, and to exploit the solved structures for the assignment of biological function to hypothetical proteins. We discuss the impact of these developments on protein classification, gene function prediction, and protein structure prediction.

Databases, Factual↗

RIKEN mouse genome encyclopedia.

We have been working to establish the comprehensive mouse full-length cDNA collection and sequence database to cover as many genes as we can, named Riken mouse genome encyclopedia. Recently we are constructing higher-level annotation (Functional ANnoTation Of Mouse cDNA; FANTOM) not only with homology search based annotation but also with expression data profile, mapping information and protein-protein database. More than 1,000,000 clones prepared from 163 tissues were end-sequenced to classify into 159,789 clusters and 60,770 representative clones were fully sequenced. As a conclusion, the 60,770 sequences contained 33,409 unique. The next generation of life science is clearly based on all of the genome information and resources. Based on our cDNA clones we developed the additional system to explore gene function. We developed cDNA microarray system to print all of these cDNA clones, protein-protein interaction screening system, protein-DNA interaction screening system and so on. The integrated database of all the information is very useful not only for analysis of gene transcriptional network and for the connection of gene to phenotype to facilitate positional candidate approach. In this talk, the prospect of the application of these genome resourced should be discussed. More information is available at the web page: http://genome.gsc.riken.go.jp/.

Animals↗

Virtual screening against metalloenzymes for inhibitors and substrates.

Molecular docking uses the three-dimensional structure of a receptor to screen databases of small molecules for potential ligands, often based on energetic complementarity. For many docking scoring functions, which calculate nonbonded interactions, metalloenzymes are challenging because of the partial covalent nature of metal-ligand interactions. To investigate how well molecular docking can identify potential ligands of metalloenzymes using a "standard" scoring function, we have docked the MDL Drug Data Report (MDDR), a functionally annotated database of 95,000 small molecules, against the X-ray crystal structures of five metalloenzymes. These enzymes included three zinc proteases, the nickel analogue of an iron enzyme, and a molybdenum metalloenzyme. The ability of the docking program to retrospectively enrich the annotated ligands as high-scoring hits for each enzyme and to calculate proper geometries was evaluated. In all five systems, the annotated ligands within the MDDR were enriched at least 20 times over random. To test the approach prospectively, a sixth target, the zinc beta-lactamase from Bacteroides fragilis, was screened against the fragment-like subset of the ZINC database. We purchased and tested 15 compounds from among the top 50 top-ranked ligands from docking, and found 5 inhibitors with apparent K(i) values less than 120 microM, the best of which was 2 microM. A more ambitious test still was predicting actual substrates for a seventh target, a Zn-dependent phosphotriesterase from Pseudomonas diminuta. Screening the Available Chemicals Directory (ACD) identified 25 thiophosphate esters as potential substrates within the top 100 ranked compounds. Eight of these, all previously uncharacterized for this enzyme, were acquired and tested, and all were confirmed experimentally as substrates. These results suggest that a simple, noncovalent scoring function may be used to identify inhibitors of at least some metalloenzymes.

Crystallography, X-Ray↗

Information decay in molecular docking screens against holo, apo, and modeled conformations of enzymes.

Molecular docking uses the three-dimensional structure of a receptor to screen a small molecule database for potential ligands. The dependence of docking screens on the conformation of the binding site remains an open question. To evaluate the information loss that occurs as the active site conformation becomes less defined, a small molecule database was docked against the holo (ligand bound), apo, and homology modeled structures of 10 different enzyme binding sites. The holo and apo representations were crystallographic structures taken from the Protein Data Bank (PDB), and the homology-modeled structures were taken from the publicly available resource ModBase. The database docked was the MDL Drug Data Report (MDDR), a functionally annotated database of 95000 small molecules that contained at least 35 ligands for each of the 10 systems. In all sites, at least 99% of the molecules in the MDDR were treated as nonbinding decoys. For each system, the holo, apo, and modeled structures were used to screen the MDDR, and the ability of each structure to enrich the known ligands for that system over random selection was evaluated. The best overall enrichment was produced by the holo structure in seven systems, the apo structure in two systems, and the modeled structure in one system. These results suggest that the performance of the docking calculation is affected by the particular representation of the receptor used in the screen, and that the holo structure is the one most likely to yield the best discrimination between known ligands and decoy molecules, but important exceptions to this rule also emerge from this study. Although each of the holo, apo, and modeled conformations led to enrichment of known ligands in all systems, the enrichment did not always rise to a level judged to be sufficient to justify the effort of a docking screen. Using a 20-fold enrichment of known ligands over random selection as a rough guideline for what might be enough to justify a docking screen, the holo conformation of the enzyme met this criterion in eight of 10 sites, whereas the apo conformation met this criterion in only two sites and the modeled conformation in three.

Animals↗

Using proteomics to mine genome sequences.

We present a method for mining unannotated or annotated genome sequences with proteomic data to identify open reading frames. The region of a genome coding for a protein sequence is identified by using information from the analysis of proteins and peptides with MALDI-TOF mass spectrometry. The raw genome sequence or any unassembled contigs of an organism are theoretically cleaved into a number of equal sized but overlapping fragments, and these are then translated in all six frames into a series of virtual proteins. Each virtual protein is then subjected to a theoretical enzymatic digestion. Standard proteomic sample preparation methods are used to separate, array, and digest the proteins of interest to peptides. The masses of the resulting peptides are measured using mass spectrometry and compared to the theoretical peptide masses of the virtual proteins. The region of the genome responsible for coding for a particular protein can then be identified when there are a large number of hits between peptides from the protein and peptides from the virtual protein. The method makes no assumptions about the location of a protein in a particular gene sequence or the positions or types of start and stop codons. To illustrate this approach, all 773 proteins of Pseudomonas aeruginosa contained in SWISS-PROT were used to theoretically test the method and optimize parameters. Increasing the size of the virtual proteins results in an overall improvement in the ability to detect the coding region, at the cost of decreasing the sensitivity of the method for smaller proteins. Increasing the minimum number of matching peptides, lowering the mass error tolerance, or increasing the signal-to-noise ratio of the simulated mass spectrum, improves the ability to detect coding regions. The method is further demonstrated on experimental data from Mycobacterium tuberculosis and is also shown to work with eukaryotic organisms (e.g., Homo sapiens).

Amino Acid Sequence↗

Differential detergent fractionation for non-electrophoretic eukaryote cell proteomics.

Differential detergent fractionation (DDF), which relies on detergents to sequentially extract proteins from eukaryotic cells, has been used to increase proteome coverage of 2D-PAGE. Here, we used DDF extraction in conjunction with the nonelectrophoretic proteomics method of liquid chromatography and electrospray ionization tandem mass spectrometry. We demonstrate that DDF can be used with 2D-LC ESI MS2 for comprehensive cellular proteomics, including a large proportion of membrane proteins. Compared to some published methods designed to isolate membrane proteins specifically, DDF extraction yields comprehensive proteomes which include twice as many membrane proteins. Two-thirds of these membrane proteins have more than one trans-membrane domain. Since DDF separates proteins based upon their physicochemistry and subcellular localization, this method also provides data useful for functional genome annotation. As more genome sequences are completed, methods which can aid in functional annotation will become increasingly important.

Animals↗

Computational approaches to protein-protein interaction.

The interactions between proteins allow the cell's life. A number of experimental, genome-wide, high-throughput studies have been devoted to the determination of protein-protein interactions and the consequent interaction networks. Here, the bioinformatics methods dealing with protein-protein interactions and interaction network are overviewed. 1. Interaction databases developed to collect and annotate this immense amount of data; 2. Automated data mining techniques developed to extract information about interactions from the published literature; 3. Computational methods to assess the experimental results developed as a consequence of the finding that the results of high-throughput methods are rather inaccurate; 4. Exploitation of the information provided by protein interaction networks in order to predict functional features of the proteins; and 5. Prediction of protein-protein interactions.

Algorithms↗

Chromosome-level genome assembly of the small-sized Taihang donkey (Equus asinus).

China harbors a rich diversity of donkey breeds, with small-sized donkeys (<110&#x2009;cm) representing a largely underexplored group. Here, we present the first high-quality, chromosome-level genome assembly of a small-sized donkey, generated using PacBio HiFi sequencing (286.7&#x2009;Gb), Hi-C scaffolding (240.47&#x2009;Gb), and annotated with RNA-seq data. The final assembly has a total length of 2.7&#x2009;Gb and comprises 32 chromosomes (including both X and Y chromosomes), in which five chromosomes were fully assembled without gaps. It possesses a scaffold N50 of 106.70&#x2009;Mb and 84 contigs (contig N50&#x2009;=&#x2009;63.60&#x2009;Mb), and captures 99.2% of BUSCO genes. The assembly achieved a consensus quality value (QV) of 77.44, corresponding to an extremely low base-level error rate, indicating exceptional nucleotide accuracy. This high-quality genome provides a valuable resource for investigating genetic variation, adaptive evolution, and domestication processes in small-sized donkeys, and will facilitate the conservation and sustainable utilization of rich donkey genetic resources in China.

Animals↗

PICRUSt2-SC: an update to the reference database used for functional prediction within PICRUSt2.

SUMMARY: PICRUSt2 is a bioinformatic tool that predicts microbial functions in amplicon sequencing data using a database of annotated reference genomes. We have constructed an updated database for PICRUSt2 that has substantially increased the number of bacterial (19,493 to 26,868) and archaeal (406 to 1,002) genomes as well as the number of functional annotations present. The previous PICRUSt2 database relied on many timely and computationally intensive manual processes that made it difficult to update. We constructed a new streamlined process to allow regular upgrades to the PICRUSt2 database on an ongoing basis, and used this process to create a new database, PICRUSt2-SC (Sugar-Coated). Additionally, we have shown that this updated database contains genomes that more closely match study sequences from a range of different environments. The genomes contained in the database therefore better represent these environments and this leads to an improvement in the predicted functional annotations obtained from PICRUSt2. AVAILABILITY AND IMPLEMENTATION: PICRUSt2 source code is freely available at https://github.com/picrust/picrust2 and at https://anaconda.org/bioconda/picrust2. The latest version of PICRUSt2 at the time of writing is also archived: https://doi.org/10.5281/zenodo.15119781. The PICRUSt2-SC database comes pre-installed with PICRUSt2 from version 2.6.0 onwards. Step-by-step instructions for making the updated database are at https://github.com/picrust/picrust2/wiki/Updating-the-PICRUSt2-database. All code used for the analyses and figures in this manuscript is at https://github.com/R-Wright-1/PICRUSt2-SC_application_note and https://doi.org/10.5281/zenodo.15119770.

Software↗

Knowledge-enhanced protein subcellular localization prediction from 3D fluorescence microscope images.

MOTIVATION: Pinpointing the subcellular location of proteins is essential for studying protein function and related diseases. Advances in spatial proteomics have shown that automatic recognition of protein subcellular localization from images could highly facilitate protein translocation analysis and biomarker discovery, but existing machine-learning works have been mostly limited to processing 2D images. By contrast, 3D images have higher spatial resolution&#xa0;and allow researchers to observe cellular structures in their natural context, but currently, there are only a few studies of 3D image processing for protein distribution analysis due to the lack of data and complexity of modeling. RESULTS: We developed a knowledge-enhanced protein subcellular localization model, KE3DLoc, which could recognize distribution patterns in 3D fluorescence microscope images using deep learning methods. The model designs an image feature extraction module that incorporates information from 3D and 2D projected cells and implements asymmetric loss and confidence weights to address data imbalance and weak cell annotation issues. Besides, considering that the biological knowledge in the Gene Ontology (GO) database can provide valuable support for protein location understanding, the KE3DLoc model incorporates a novel knowledge enhancement module that optimizes the protein representation by related knowledge graphs derived from the GO. Since the image module and the knowledge module calculate features from different levels, KE3DLoc designs protein ID aggregation to enhance the consistency of protein features across different cells. Experimental results on three public datasets have demonstrated that the KE3DLoc significantly outperforms existing methods and provides valuable insights for spatial proteomics research. AVAILABILITY AND IMPLEMENTATION: All datasets and codes used in this study are available at GitHub: https://github.com/PRBioimages/KE3DLoc.

Microscopy, Fluorescence↗

BioEditor-simplifying macromolecular structure annotation.

SUMMARY: BioEditor is an application to enable scientists and educators to prepare and present structure annotations containing formatted text, graphics, sequence data, and interactive molecular views. It is intended to bridge the gap between printed journal articles and Internet presentation formats. BioEditor is relevant in the era of structural genomics, where annotation and publication could become the rate determining step in structure determination. AVAILABILITY: BioEditor is available at http://bioeditor.sdsc.edu. The Web site includes the latest version of the software for Microsoft Windows, including documentation, the opportunity to submit bug reports and suggestions, example documentaries prepared with BioEditor and a repository where users can submit documentaries for posting to the site.

Biopolymers↗

Genes in canine articular cartilage that respond to mechanical injury: gene expression studies with Affymetrix canine GeneChip.

The Affymetrix canine GeneChip with 23,836 probe sets was used to look for cartilage genes that are significantly altered in response to mechanical impact. The model using canine articular cartilage explants loaded in vitro has been described previously (Chen et al., J Orthop Res 19:703-711, 2001). It is our hypothesis that genes that are activated or repressed in articular cartilage after impact injury initiate cartilage degeneration, leading to osteoarthritis in dogs. Gene expression of known cartilage genes was generally consistent with cartilage biology. A total of 528 genes were significantly (P < .01) up- or down- regulated in response to mechanical damage. After applying the strict Bonferroni correction, 172 remained significantly affected. One of these genes, MIG-6/gene 33, was chosen for verification by real- time quantitative reverse transcriptase polymerase chain reaction (RT-PCR). A 3.8- fold increase in expression was confirmed, consistent with the microarray chip data. Deficiencies in the current annotation of the canine chip are discussed. Gene expression studies with the Affymetrix canine GeneChip are potentially valuable, but await more complete annotation.

Animals↗

Automating the identification of DNA variations using quality-based fluorescence re-sequencing: analysis of the human mitochondrial genome.

Diagnostic re-sequencing plays a central role in medical and evolutionary genetics. In this report we describe a process that applies fluorescence-based re-sequencing and an integrated set of analysis tools to automate and simplify the identification of DNA variations using the human mitochondrial genome as a model system. Two programs used in genome sequence analysis (Phred, a base-caller, and Phrap, a sequence assembler) are applied to assess the quality of each base call across the sequence. Potential DNA variants are automatically identified and 'tagged' by comparing the assembled sequence with a reference sequence. We also show that employing the Consed program to display a set of highly annotated reference sequences greatly simplifies data analysis by providing a visual database containing information on the location of the PCR primers, coding and regulatory sequences and previously known DNA variants. Among the 12 genomes sequenced 378 variants including 29 new variants were identified along with two heteroplasmic sites, automatically detected by the PolyPhred program. Overall we document the ease and speed of performing high quality and accurate fluorescence-based re-sequencing on long tracts of DNA as well as the application of new approaches to automatically find and view DNA variants among these sequences.

Base Sequence↗

CHOP: parsing proteins into structural domains.

Sequence-based domain assignment is one of the most important and challenging problems in structural biology. We have developed a method, CHOP, that chops proteins into domain-like fragments. The basic idea is to cut proteins from entirely sequenced organisms beginning from very reliable experimental information (Protein Data Bank), proceeding to expert annotations of domain-like regions (Pfam-A) and completing through cuts based on termini of native protein ends. The CHOP server takes protein sequences as input and returns the dissections supported by homology transfer. CHOP results are precompiled for many entirely sequenced proteomes. The service is available at http://www.rostlab.org/services/CHOP/.

Algorithms↗

TFBScluster web server for the identification of mammalian composite regulatory elements.

Identification of transcriptional regulatory elements represents a critical step in our ability to reconstruct transcriptional regulatory networks from gene expression profiling datasets. To facilitate computational identification of candidate gene regulatory elements from whole genome sequences, we have developed the TFBScluster web server that integrates several tools for the genome-wide identification and subsequent characterization of transcription factor binding site clusters that are conserved in multiple mammalian species. Either the human or mouse genomes can be used as the reference sequence with direct links from the search results to the ENSEMBL and UCSC genome browsers. Moreover, TFBScluster provides seamless integration of transcription factor binding site searches with genome annotation and gene expression profiling data, to allow prioritising computational predictions for subsequent experimental validation. TFBScluster is publicly available at http://hscl.cimr.cam.ac.uk/TFBScluster_genome_portal.html.

Animals↗

Global survey of chromatin accessibility using DNA microarrays.

An increasing number of studies indicate a central role for chromatin remodeling in the regulation of gene expression. Current methods for high-resolution studies of the relationship between chromatin accessibility and transcription are low throughput, making a genome-wide study impractical. To enable the simultaneous measurement of the global chromatin accessibility state at the resolution of single genes, we developed the Chromatin Array technique, in which chromatin is separated by its condensation state using either the solubility differences of mono- and oligonucleosomes in specific buffers or controlled DNase I digestion and selection of the large refractory (condensed) DNA fragments. By probing with a comparative genomic hybridization style microarray, we can determine the condensation state of thousands of individual loci and correlate this with transcriptional activity. Applying this technique to the breast tumor model cell line, MCF7, we found that when the condensation is homogeneous in the population of cells, expression is inversely proportional to the level of accessibility and the two methods of accessibility-based target selection correlate well. Using functional annotation and comparative genomic hybridization data, we have begun to decipher the possible biological implications of the relationship between chromatin accessibility and expression.

Breast Neoplasms↗

Solution structure of Archaeglobus fulgidis peptidyl-tRNA hydrolase (Pth2) provides evidence for an extensive conserved family of Pth2 enzymes in archea, bacteria, and eukaryotes.

The solution structure of protein AF2095 from the thermophilic archaea Archaeglobus fulgidis, a 123-residue (13.6-kDa) protein, has been determined by NMR methods. The structure of AF2095 is comprised of four alpha-helices and a mixed beta-sheet consisting of four parallel and anti-parallel beta-strands, where the alpha-helices sandwich the beta-sheet. Sequence and structural comparison of AF2095 with proteins from Homo sapiens, Methanocaldococcus jannaschii, and Sulfolobus solfataricus reveals that AF2095 is a peptidyl-tRNA hydrolase (Pth2). This structural comparison also identifies putative catalytic residues and a tRNA interaction region for AF2095. The structure of AF2095 is also similar to the structure of protein TA0108 from archaea Thermoplasma acidophilum, which is deposited in the Protein Data Bank but not functionally annotated. The NMR structure of AF2095 has been further leveraged to obtain good-quality structural models for 55 other proteins. Although earlier studies have proposed that the Pth2 protein family is restricted to archeal and eukaryotic organisms, the similarity of the AF2095 structure to human Pth2, the conservation of key active-site residues, and the good quality of the resulting homology models demonstrate a large family of homologous Pth2 proteins that are conserved in eukaryotic, archaeal, and bacterial organisms, providing novel insights in the evolution of the Pth and Pth2 enzyme families.

Archaea↗