Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44Linked to original sources

An agent-based system for re-annotation of genomes.

Genome annotation projects can produce incorrect results if they are based on obsolete data or inappropriate models. We have developed an automatic re-annotation system that uses agents to perform repetitive tasks and reports the results to the user. These tasks involve BLAST searches on biological databases (GenBank) and the use of detection tools (Genemark and Glimmer) to identify new open reading frames. Several agents execute these tools and combine their results to produce a list of open reading frames that is sent back to the user. Our goal was to reduce the manual work, executing most tasks automatically by computational tools. A prototype was implemented and validated using Mycoplasma pneumoniae and Haemophilus influenzae original annotated genomes. The results reported by the system identify most of new features present in the re-annotated versions of these genomes.

Computational Biology↗

Structuring the universe of proteins.

High-throughput sequencing of human genomes and those of important model organisms (mouse, Drosophila melanogaster, Caenorhabditis elegans, fungi, archaea) and bacterial pathogens has laid the foundation for another "big science" initiative in biology. Together, X-ray crystallographers, nuclear magnetic resonance (NMR) spectroscopists, and computational biologists are pursuing high-throughput structural studies aimed at developing a comprehensive three-dimensional view of the protein structure universe. The new science of structural genomics promises more than 10,000 experimental protein structures and millions of calculated homology models of related proteins. The evolutionary underpinnings and technological challenges of automating target selection, protein expression and purification, sample preparation, NMR and X-ray data measurement/analysis, homology modeling, and structure/function annotation are discussed in detail. An informative case study from one of the structural genomics centers funded by the National Institutes of Health and the National Institute of General Medical Sciences (NIH/NIGMS) demonstrates how this experimental/computational pipeline will reveal important links between form and function in biology and provide new insights into evolution and human health and disease.

Amino Acid Sequence↗

Clinical anthropometry and medical genetics: a compilation of body measurements in genetic and congenital disorders.

Anthropometry has become an important tool in the study of genetic conditions, particularly as a diagnostic aid for the clinical geneticist. However, many practicing physicians do not do anthropometry of patients for several reasons, such as: appropriate measurements in a given situation are unknown; normative reference data are unavailable; or analysis and interpretation of the data are confusing. In this review we present an annotated compilation of informative measurements for hereditary and congenital disorders and a guide to normative anthropometric data of use in evaluation and diagnosis of such disorders. Further development of multivariate approaches will enhance the application of anthropometry as a means of identifying and classifying a syndrome and documenting the natural history of many disorders. Continued cooperation among physicians, geneticists, and anthropologists for the collection and assessment of patient and normative data is essential if these goals are to be realized.

Anthropometry↗

Evaluation of annotation strategies using an entire genome sequence.

MOTIVATION: Genome-wide functional annotation either by manual or automatic means has raised considerable concerns regarding the accuracy of assignments and the reproducibility of methodologies. In addition, a performance evaluation of automated systems that attempt to tackle sequence analyses rapidly and reproducibly is generally missing. In order to quantify the accuracy and reproducibility of function assignments on a genome-wide scale, we have re-annotated the entire genome sequence of Chlamydia trachomatis (serovar D), in a collaborative manner. RESULTS: We have encoded all annotations in a structured format to allow further comparison and data exchange and have used a scale that records the different levels of potential annotation errors according to their propensity to propagate in the database due to transitive function assignments. We conclude that genome annotation may entail a considerable amount of errors, ranging from simple typographical errors to complex sequence analysis problems. The most surprising result of this comparative study is that automatic systems might perform as well as the teams of experts annotating genome sequences.

Amino Acid Sequence↗

Scalable approaches for functional analyses of whole-genome sequencing non-coding variants.

Non-coding genetic variants outside of protein-coding genome regions play an important role in genetic and epigenetic regulation. It has become increasingly important to understand their roles, as non-coding variants often make up the majority of top findings of genome-wide association studies (GWAS). In addition, the growing popularity of disease-specific whole-genome sequencing (WGS) efforts expands the library of and offers unique opportunities for investigating both common and rare non-coding variants, which are typically not detected in more limited GWAS approaches. However, the sheer size and breadth of WGS data introduce additional challenges to predicting functional impacts in terms of data analysis and interpretation. This review focuses on the recent approaches developed for efficient, at-scale annotation and prioritization of non-coding variants uncovered in WGS analyses. In particular, we review the latest scalable annotation tools, databases and functional genomic resources for interpreting the variant findings from WGS based on both experimental data and in silico predictive annotations. We also review machine learning-based predictive models for variant scoring and prioritization. We conclude with a discussion of future research directions which will enhance the data and tools necessary for the effective functional analyses of variants identified by WGS to improve our understanding of disease etiology.

Genome-Wide Association Study↗

Genome2D: a visualization tool for the rapid analysis of bacterial transcriptome data.

Genome2D is a Windows-based software tool for visualization of bacterial transcriptome and customized datasets on linear chromosome maps constructed from annotated genome sequences. Genome2D facilitates the analysis of transcriptome data by using different color ranges to depict differences in gene-expression levels on a genome map. Such output format enables visual inspection of the transcriptome data, and will quickly reveal transcriptional units, without prior knowledge of expression level cutoff values. The compiled version of Genome2D is freely available for academic or non-profit use from http://molgen.biol.rug.nl/molgen/research/molgensoftware.php.

Bacterial Proteins↗

GeneLook: a novel ab initio gene identification system suitable for automated annotation of prokaryotic sequences.

With the rapid increases in the amounts of sequence data for prokaryotic genomes, it has become important to develop systems for automated and accurate genome annotation. We present herein a novel ab initio gene identification system, GeneLook, that predicts protein-coding open reading frames (ORFs) with high sensitivity and specificity with no prior knowledge of the sequence composition. The system predicts protein-coding ORFs in two stages, seed ORF selection and main prediction. In the selection of reliable seed ORFs containing at least 200 codons, GeneLook predicts translation start sites and operon structures through searches for ribosome-binding sites and a novel operon prediction algorithm. The codon and nucleotide frequencies of seed ORFs are then used to determine values for two new coding-potential parameters for identification of protein-coding ORFs of at least 34 codons and for another parameter that improves the prediction accuracy for GC-rich genomes. In the main prediction, GeneLook uses these parameters to identify the most likely genes of a given minimal length. We assessed the performance of GeneLook with two indices, sensitivity and specificity that are defined as true positives (TP)/(TP+false negatives) and TP/(TP+false positives), respectively. This system predicted protein-coding ORFs for Escherichia coli and Bacillus subtilis with sensitivities of 96.5% and 96.2%, respectively, and specificities of 96.9% and 96.1%, respectively. The system also identified 94.1% of annotated genes of the Pseudomonas aeruginosa genome, which is GC-rich, with high specificity (97.2%). Furthermore, GeneLook identified protein-coding ORFs with high accuracy from a wide variety of prokaryotic genomes.

Bacillus subtilis↗

Integrated ¹H-NMR Metabolomics and Growth Kinetics Uncover Three Distinct Metabolic Scenarios in Lactiplantibacillus pentosus P7 Fermentation of Plant-Derived Prebiotics.

Lactic acid bacteria (LAB) drive a broad range of food and biotechnological fermentations, the outcomes of which depend not only on the bacterial genotype but also on the chemical composition of the fermentation substrate. To resolve how a single strain reorganises chemically distinct plant matrices, we profiled fermentations of Lactiplantibacillus pentosus P7 (GenBank JBLMKZ000000000) on garlic, onion, and kiwifruit extracts prepared in water and 70% ethanol, using growth kinetics combined with solvent-suppressed 500-MHz proton nuclear magnetic resonance metabolomics over 48 h, and integrated the data with whole-genome pathway annotations. Three substrate-specific metabolic scenarios emerged. On garlic, P7 grew vigorously, with the water extract exceeding the de Man-Rogosa-Sharpe reference medium at every time point (peak ΔOD₆₀₀ of 9.38 versus 8.52 at 24 h) and accumulating sorbose, rhamnose, and the aromatic amino acids phenylalanine and tryptophan (3.04- to 3.70-fold increases), providing first metabolic evidence consistent with the strain's four-copy aroE shikimate-dehydrogenase expansion. On onion, the lowest cell density coincided with the highest lactate output of the dataset (5.21-fold rise at 48 h), transient 5-hydroxymethylfurfural reduction, and accumulation of acetoin and 1,3-propanediol, mapping onto a redundant set of pyridine-nucleotide-dependent oxidoreductases and a pdu-independent diol pathway. On kiwifruit, citrate accumulated 8.9-fold at 16 h and then declined, consistent with an intact citCDEFG citrate-lyase operon paired with absence of canonical oxidative tricarboxylic acid enzymes. The optimal extraction solvent was substrate-dependent, water for garlic and ethanol for onion and kiwifruit. Overall, these results show that substrate chemistry, rather than strain identity, dictates which genome-encoded pathways P7 engages, establishing P7 as a versatile, substrate-tunable platform for the functional fermentation and biorefining of furanic-rich substrate streams. Raw NMR data and ISA-Tab metadata are available via MetaboLights with identifier MTBLS14463.

Lactiplantibacillus pentosus↗

The use of MPSS for whole-genome transcriptional analysis in Arabidopsis.

We have generated 36,991,173 17-base sequence "signatures" representing transcripts from the model plant Arabidopsis. These data were derived by massively parallel signature sequencing (MPSS) from 14 libraries and comprised 268,132 distinct sequences. Comparable data were also obtained with 20-base signatures. We developed a method for handling these data and for comparing these signatures to the annotated Arabidopsis genome. As part of this procedure, 858,019 potential or "genomic" signatures were extracted from the Arabidopsis genome and classified based on the position and orientation of the signatures relative to annotated genes. A comparison of genomic and expressed signatures matched 67,735 signatures predicted to be derived from distinct transcripts and expressed at significant levels. Expressed signatures were derived from the sense strand of at least 19,088 of 29,084 annotated genes. A comparison of the genomic and expression signatures demonstrated that approximately 7.7% of genomic signatures were underrepresented in the expression data. These genomic signatures contained one of 20 four-base words that were consistently associated with reduced MPSS abundances. More than 89% of the sum of the expressed signature abundances matched the Arabidopsis genome, and many of the unmatched signatures found in high abundances were predicted to match to previously uncharacterized transcripts.

Arabidopsis↗

Gene and alternative splicing annotation with AIR.

Designing effective and accurate tools for identifying the functional and structural elements in a genome remains at the frontier of genome annotation owing to incompleteness and inaccuracy of the data, limitations in the computational models, and shifting paradigms in genomics, such as alternative splicing. We present a methodology for the automated annotation of genes and their alternatively spliced mRNA transcripts based on existing cDNA and protein sequence evidence from the same species or projected from a related species using syntenic mapping information. At the core of the method is the splice graph, a compact representation of a gene, its exons, introns, and alternatively spliced isoforms. The putative transcripts are enumerated from the graph and assigned confidence scores based on the strength of sequence evidence, and a subset of the high-scoring candidates are selected and promoted into the annotation. The method is highly selective, eliminating the unlikely candidates while retaining 98% of the high-quality mRNA evidence in well-formed transcripts, and produces annotation that is measurably more accurate than some evidence-based gene sets. The process is fast, accurate, and fully automated, and combines the traditionally distinct gene annotation and alternative splicing detection processes in a comprehensive and systematic way, thus considerably aiding in the ensuing manual curation efforts.

Alternative Splicing↗

An HMM model for coiled-coil domains and a comparison with PSSM-based predictions.

MOTIVATION: Large-scale sequence data require methods for the automated annotation of protein domains. Many of the predictive methods are based either on a Position Specific Scoring Matrix (PSSM) of fixed length or on a window-less Hidden Markov Model (HMM). The performance of the two approaches is tested for Coiled-Coil Domains (CCDs). The prediction of CCDs is used frequently, and its optimization seems worthwhile. RESULTS: We have conceived MARCOIL, an HMM for the recognition of proteins with a CCD on a genomic scale. A cross-validated study suggests that MARCOIL improves predictions compared to the traditional PSSM algorithm, especially for some protein families and for short CCDs. The study was designed to reveal differences inherent in the two methods. Potential confounding factors such as differences in the dimension of parameter space and in the parameter values were avoided by using the same amino acid propensities and by keeping the transition probabilities of the HMM constant during cross-validation. AVAILABILTY: The prediction program and the databases are available at http://www.wehi.edu.au/bioweb/Mauro/Marcoil

Algorithms↗

The complete genome and proteome of Mycoplasma mobile.

Although often considered "minimal" organisms, mycoplasmas show a wide range of diversity with respect to host environment, phenotypic traits, and pathogenicity. Here we report the complete genomic sequence and proteogenomic map for the piscine mycoplasma Mycoplasma mobile, noted for its robust gliding motility. For the first time, proteomic data are used in the primary annotation of a new genome, providing validation of expression for many of the predicted proteins. Several novel features were discovered including a long repeating unit of DNA of approximately 2435 bp present in five complete copies that are shown to code for nearly identical yet uniquely expressed proteins. M. mobile has among the lowest DNA GC contents (24.9%) and most reduced set of tRNAs of any organism yet reported (28). Numerous instances of tandem duplication as well as lateral gene transfer are evident in the genome. The multiple available complete genome sequences for other motile and immotile mycoplasmas enabled us to use comparative genomic and phylogenetic methods to suggest several candidate genes that might be involved in motility. The results of these analyses leave open the possibility that gliding motility might have arisen independently more than once in the mycoplasma lineage.

Amino Acid Sequence↗

Bioinformatics issues for automating the annotation of genomic sequences.

The rapid explosion in the amount of biological data being generated worldwide is surpassing efforts to manage analysis of the data. As part of an ongoing project to automate and manage bioinformatics analysis, the authors have designed and implemented a simple automated annotation system, which is described in this paper. The system is applied to existing GenBank/DDBJ/EMBL entries and compared with existing annotations to illustrate not only potential errors but also that they are generally not up-to-date, as a result of new versions of analysis tools and updates of genomic repositories. We highlight the important Bioinformatics issues of storage and management of information to ensure data and results are kept up-to-date in light of new information becoming available. Surprisingly, from just four database entries, a significant number of new features were found. We describe the results as well as identify important issues that need to be addressed in order to automate the re-analysis/re-annotation of genomic sequences within a reasonable timeframe.

Computational Biology↗

Microarray analysis of orthologous genes: conservation of the translational machinery across species at the sequence and expression level.

BACKGROUND: Genome projects have provided a vast amount of sequence information. Sequence comparison between species helps to establish functional catalogues within organisms and to study how they are maintained and modified across phylogenetic groups during evolution. Microarray studies allow us to determine groups of genes with similar temporal regulation and perhaps also common regulatory upstream regions for binding of transcription factors. The integration of sequence and expression data is expected to refine our current annotations and provide some insight into the evolution of gene regulation across organisms. RESULTS: We have investigated how well the protein subcellular localization and functional categories established from clustering of orthologous genes agree with gene-expression data in Saccharomyces cerevisiae. An increase in the resolution of biologically meaningful classes is observed upon the combination of experiments under different conditions. The functional categories deduced by sequence comparison approaches are, in general, preserved at the level of expression and can sometimes interact into larger co-regulated networks, such as the protein translation process. Differences and similarities in the expression between cytoplasmic-mitochondrial and interspecies translation machineries complement evolutionary information from sequence similarity. CONCLUSIONS: Combination of several microarray experiments is a powerful tool for the identification of upstream regulatory motifs of yeast genes involved in protein synthesis. Comparison of these yeast co-regulated genes against the archaeal and bacterial operons indicates that the components of the protein translation process are conserved across organisms at the expression level with minor specific adaptations.

Archaea↗

Cognitive computer-based video analysis: its application in assessing the usability of medical systems.

This paper describes a methodology, based on cognitive research, for assessing the usability of medical computing systems. The issue of developing appropriate evaluation tools, both for use in the design process and for analysis of end products, is beginning to be recognized as being of great importance. In this paper, the use of video recording for collecting empirical data on system usability is detailed. The techniques described allow for the collection of an integrated data set consisting of the transcripts of physicians as they "think aloud" in interacting with a medical system, along with video records of user-computer interaction. The use of coding methodologies and a computer-based annotation system for the analysis of video data are described. Our preliminary experience indicates that this methodology offers a powerful way for assessing physicians' informational needs. Implications for the development and evaluation of medical information systems are discussed.

Computer Systems↗

Reference genome of the Californian trapdoor spider Aptostichus stephencolberti Bond 2008 (Araneae: Mygalomorphae: Euctenizidae).

We present a reference genome assembly for the trapdoor spider Aptostichus stephencolberti. This species, described in 2008, is endemic to the highly fragmented coastal dune habitats of Northern California from Monterey to the San Francisco Bay Area. Trapdoor spiders are ideal taxa for landscape scale genomic studies owing to their extreme site fidelity and limited dispersal capabilities; these same characteristics make them prone to extinction. Genomic studies of species like A. stephencolberti can reveal novel areas of endemism and high conservation value that may not be evident in species with wider ranges and greater dispersal capabilities. As part of the California Conservation Genomics Project, we constructed the A. stephencolberti reference genome from high quality long-read sequences, scaffolded with proximity ligation Omni-C data. The primary assembly comprises 551 scaffolds spanning 3.63 Gbp, a scaffold N50 of 62.2 Mbp and BUSCO completeness of 95.6%. We estimate 52 chromosomes yet find no (TTAGG)n telomer repeats. Expanding the telomeric repeat search finds an ancestral loss of the repeat from all spiders. Automated annotation using the NCBI refseq pipeline and RNAseq data from whole adults finds 14,067 genes with a BUSCO annotation completeness of 95.56%. Repeat annotation identified 77% of the genome to be interspersed repeats. This resource, the first for family Euctenizidae will facilitate future study and resulting conservation actions of A. stephencolberti and other Aptostichus sp. populations associated with the rapidly changing California coastal dune ecosystem.

Aptostichus stephencolberti↗

A complexity reduction algorithm for analysis and annotation of large genomic sequences.

DNA is a universal language encrypted with biological instruction for life. In higher organisms, the genetic information is preserved predominantly in an organized exon/intron structure. When a gene is expressed, the exons are spliced together to form the transcript for protein synthesis. We have developed a complexity reduction algorithm for sequence analysis (CRASA) that enables direct alignment of cDNA sequences to the genome. This method features a progressive data structure in hierarchical orders to facilitate a fast and efficient search mechanism. CRASA implementation was tested with already annotated genomic sequences in two benchmark data sets and compared with 15 annotation programs (10 ab initio and 5 homology-based approaches) against the EST database. By the use of layered noise filters, the complexity of CRASA-matched data was reduced exponentially. The results from the benchmark tests showed that CRASA annotation excelled in both the sensitivity and specificity categories. When CRASA was applied to the analysis of human Chromosomes 21 and 22, an additional 83 potential genes were identified. With its large-scale processing capability, CRASA can be used as a robust tool for genome annotation with high accuracy by matching the EST sequences precisely to the genomic sequences.

Algorithms↗

'The 39 steps' in gene expression profiling: critical issues and proposed best practices for microarray experiments.

Gene expression microarrays have been used widely to address increasingly complex biological questions and to produce an unprecedented amount of data, but have yet to realize their full potential. The interpretation of microarray data remains a major challenge because of the complexity of the underlying biological networks. To gather meaningful expression data, it is crucial to develop standardized approaches for vigilant study design, controlled annotation of resources, careful quality control of experiments, robust statistics, and data registration and storage. This article reviews the steps needed in the design and execution of valid microarray experiments so that global gene expression data can play a major role in the pursuit of future biological discoveries that will impact drug development.

Drug Design↗