Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “genome annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

theBIGbam: compression and interactive exploration of large-scale sequencing alignments with circular mapping support.

SUMMARY: theBIGbam (github.com/bhagavadgitadu22/theBIGbam) is a genome browser and alignment viewer designed for massive metagenomic and metatranscriptomic datasets. The tool takes BAM files containing read alignments, together with genome assemblies in FASTA format or annotated genome sequences in GenBank format. Alternatively, it can start from raw FASTQ reads and generate alignments using a modified mapper that supports circular genomes, enabling seamless read mapping across genome ends. theBIGbam can compress hundreds of gigabytes of input files 10- to 100-fold into dedicated databases while retaining key per-position information, including coverage depth and recurrent mismatches, insertions, and deletions between reads and the reference. These databases can be served to a local web browser, enabling interactive exploration of any contig in any sample using DNAFeaturesViewer for genome maps and Bokeh for mapping-derived features. Contig-sample pairs available for visualization can be filtered using a range of summary metrics calculated per contig, per sample, and per contig-sample pair to guide users toward the most relevant signals. Through its interactive visualization, theBIGbam facilitates the exploration of complex datasets, while its integrated database-combining assembly features, annotated features, and mapping-derived features-provides the information needed to investigate biological hypotheses systematically. Designed to complement existing browsing tools like IGV and Anvi'o, theBIGbam is particularly suited for examining misassemblies, subpopulations, microdiversity, and contig topology in large-scale datasets. AVAILABILITY AND IMPLEMENTATION: theBIGbam is an open-source Rust/Python package that can be installed from Bioconda or PyPI. The source code and documentation are available on GitHub (github.com/bhagavadgitadu22/theBIGbam).

Software↗

ZCURVE_V: a new self-training system for recognizing protein-coding genes in viral and phage genomes.

BACKGROUND: It necessary to use highly accurate and statistics-based systems for viral and phage genome annotations. The GeneMark systems for gene-finding in virus and phage genomes suffer from some basic drawbacks. This paper puts forward an alternative approach for viral and phage gene-finding to improve the quality of annotations, particularly for newly sequenced genomes. RESULTS: The new system ZCURVE_V has been run for 979 viral and 212 phage genomes, respectively, and satisfactory results are obtained. To have a fair comparison with the currently available software of similar function, GeneMark, a total of 30 viral genomes that have not been annotated by GeneMark are selected to be tested. Consequently, the average specificity of both systems is well matched, however the average sensitivity of ZCURVE_V for smaller viral genomes (< 100 kb), which constitute the main parts of viral genomes sequenced so far, is higher than that of GeneMark. Additionally, for the genome of Amsacta moorei entomopoxvirus, probably with the lowest genomic GC content among the sequenced organisms, the accuracy of ZCURVE_V is much better than that of GeneMark, because the later predicts hundreds of false-positive genes. ZCURVE_V is also used to analyze well-studied genomes, such as HIV-1, HBV and SARS-CoV. Accordingly, the performance of ZCURVE_V is generally better than that of GeneMark. Finally, ZCURVE_V may be downloaded and run locally, particularly facilitating its utilization, whereas GeneMark is not downloadable. Based on the above comparison, it is suggested that ZCURVE_V may serve as a preferred gene-finding tool for viral and phage genomes newly sequenced. However, it is also shown that the joint application of both systems, ZCURVE_V and GeneMark, leads to better gene-finding results. The system ZCURVE_V is freely available at: http://tubic.tju.edu.cn/Zcurve_V/. CONCLUSION: ZCURVE_V may serve as a preferred gene-finding tool used for viral and phage genomes, especially for anonymous viral and phage genomes newly sequenced.

Algorithms↗

A gene-based high-resolution comparative radiation hybrid map as a framework for genome sequence assembly of a bovine chromosome 6 region associated with QTL for growth, body composition, and milk performance traits.

BACKGROUND: A number of different quantitative trait loci (QTL) for various phenotypic traits, including milk production, functional, and conformation traits in dairy cattle as well as growth and body composition traits in meat cattle, have been mapped consistently in the middle region of bovine chromosome 6 (BTA6). Dense genetic and physical maps and, ultimately, a fully annotated genome sequence as well as their mutual connections are required to efficiently identify genes and gene variants responsible for genetic variation of phenotypic traits. A comprehensive high-resolution gene-rich map linking densely spaced bovine markers and genes to the annotated human genome sequence is required as a framework to facilitate this approach for the region on BTA6 carrying the QTL. RESULTS: Therefore, we constructed a high-resolution radiation hybrid (RH) map for the QTL containing chromosomal region of BTA6. This new RH map with a total of 234 loci including 115 genes and ESTs displays a substantial increase in loci density compared to existing physical BTA6 maps. Screening the available bovine genome sequence resources, a total of 73 loci could be assigned to sequence contigs, which were already identified as specific for BTA6. For 43 loci, corresponding sequence contigs, which were not yet placed on the bovine genome assembly, were identified. In addition, the improved potential of this high-resolution RH map for BTA6 with respect to comparative mapping was demonstrated. Mapping a large number of genes on BTA6 and cross-referencing them with map locations in corresponding syntenic multi-species chromosome segments (human, mouse, rat, dog, chicken) achieved a refined accurate alignment of conserved segments and evolutionary breakpoints across the species included. CONCLUSION: The gene-anchored high-resolution RH map (1 locus/300 kb) for the targeted region of BTA6 presented here will provide a valuable platform to guide high-quality assembling and annotation of the currently existing bovine genome sequence draft to establish the final architecture of BTA6. Hence, a sequence-based map will provide a key resource to facilitate prospective continued efforts for the selection and validation of relevant positional and functional candidates underlying QTL for milk production and growth-related traits mapped on BTA6 and on similar chromosomal regions from evolutionary closely related species like sheep and goat. Furthermore, the high-resolution sequence-referenced BTA6 map will enable precise identification of multi-species conserved chromosome segments and evolutionary breakpoints in mammalian phylogenetic studies.

Animals↗

Proteomic analysis using an unfinished bacterial genome: the effects of subminimum inhibitory concentrations of antibiotics on Mannheimia haemolytica virulence factor expression.

Here we identify, using nonelectrophoretic proteomics, effects of subminimum inhibitory concentrations (subMIC) of two antibiotic preparations, chlortetracycline (CTC), and chlortetracycline-sulfamethazine (CTC + SMZ), on protein expression in the bovine respiratory pathogen Mannheimia haemolytica. The M. haemolytica genome is currently in draft form, and annotation is incomplete. Relying on the principle of gene sequence conservation across species, we used annotated genomes from closely related species to identify, confirm, and functionally annotate 495 M. haemolytica proteins. To conduct quantitative comparative proteomics, we developed a protein quantitation method based on the cross correlation function of the SEQUEST algorithm. When M. haemolytica was cultivated in the presence of 1/4 MIC of CTC and CTC + SMZ, expression of proteins involved in energy production, nucleotide metabolism, translation, and the bacterial stress response (chaperones) were affected. The most notable subMIC effect was a significant decrease in the expression of leukotoxin A, which is an important M. haemolytica virulence factor. Reduction in leukotoxin expression could be one of the molecular mechanisms responsible for the efficacy of these antibiotics against bovine respiratory disease.

Algorithms↗

Genomic approaches to the genetics of alcoholism.

When studying complex diseases such as alcoholism that develop as a result of numerous genetic and environmental factors, researchers can use the sequence data that have become available both for the human and for animal genomes. For these analyses, investigators are being aided by efforts to identify and characterize functionally relevant DNA sequences in the entire genomic DNA sequence--a process called annotation. Various bioinformatics and annotation tools can help in this enterprise. These include four primary approaches: (1) precomputed, annotated public Web sites that provide a plethora of information; (2) in-house analyses from which users can choose the appropriate analyses for their purposes; (3) Web-based annotation systems that analyze a user's DNA sequence; and (4) private resources that provide access to annotated genomic sequences at cost. In addition to careful study of the DNA sequence for clues about function, expression studies of mRNA levels using gene chips provide information about the activity levels of thousands of genes that may vary in different tissues, different animals and people, or under different environmental conditions.

Alcoholism↗

GeneFarm, structural and functional annotation of Arabidopsis gene and protein families by a network of experts.

Genomic projects heavily depend on genome annotations and are limited by the current deficiencies in the published predictions of gene structure and function. It follows that, improved annotation will allow better data mining of genomes, and more secure planning and design of experiments. The purpose of the GeneFarm project is to obtain homogeneous, reliable, documented and traceable annotations for Arabidopsis nuclear genes and gene products, and to enter them into an added-value database. This re-annotation project is being performed exhaustively on every member of each gene family. Performing a family-wide annotation makes the task easier and more efficient than a gene-by-gene approach since many features obtained for one gene can be extrapolated to some or all the other genes of a family. A complete annotation procedure based on the most efficient prediction tools available is being used by 16 partner laboratories, each contributing annotated families from its field of expertise. A database, named GeneFarm, and an associated user-friendly interface to query the annotations have been developed. More than 3000 genes distributed over 300 families have been annotated and are available at http://genoplante-info.infobiogen.fr/Genefarm/. Furthermore, collaboration with the Swiss Institute of Bioinformatics is underway to integrate the GeneFarm data into the protein knowledgebase Swiss-Prot.

Arabidopsis↗

miRGen: a database for the study of animal microRNA genomic organization and function.

miRGen is an integrated database of (i) positional relationships between animal miRNAs and genomic annotation sets and (ii) animal miRNA targets according to combinations of widely used target prediction programs. A major goal of the database is the study of the relationship between miRNA genomic organization and miRNA function. This is made possible by three integrated and user friendly interfaces. The Genomics interface allows the user to explore where whole-genome collections of miRNAs are located with respect to UCSC genome browser annotation sets such as Known Genes, Refseq Genes, Genscan predicted genes, CpG islands and pseudogenes. These miRNAs are connected through the Targets interface to their experimentally supported target genes from TarBase, as well as computationally predicted target genes from optimized intersections and unions of several widely used mammalian target prediction programs. Finally, the Clusters interface provides predicted miRNA clusters at any given inter-miRNA distance and provides specific functional information on the targets of miRNAs within each cluster. All of these unique features of miRGen are designed to facilitate investigations into miRNA genomic organization, co-transcription and targeting. miRGen can be freely accessed at http://www.diana.pcbi.upenn.edu/miRGen.

Animals↗

How to interpret an anonymous bacterial genome: machine learning approach to gene identification.

In this report we address the problem of accurate statistical modeling of DNA sequences, either coding or noncoding, for a bacterial species whose genome (or a large portion) was sequenced but not yet characterized experimentally. Availability of these models is critical for successful solution of the genome annotation task by statistical methods of gene finding. We present the method, GeneMark-Genesis, which learns the parameters of Markov models of protein-coding and noncoding regions from anonymous bacterial genomic sequence. These models are subsequently used in the GeneMark and GeneMark.hmm gene-finding programs. Although there is basically one model of a noncoding region for a given genome, several models of protein-coding region are automatically obtained by GeneMark-Genesis. The diversity of protein-coding models reflects the diversity of oligonucleotide compositions, particularly the diversity of codon usage strategies observed in genes from one and the same genome. In the simplest and the most important case, there are just two gene models-typical and atypical ones. We show that the atypical model allows one to predict genes that escape identification by the typical model. Many genes predicted by the atypical model appear to be horizontally transferred genes. The early versions of GeneMark-Genesis were used for annotating the genomes of Methanoccocus jannaschii and Helicobacter pylori. We report the results of accuracy testing of the full-scale version of GeneMark-Genesis on 10 completely sequenced bacterial genomes. Interestingly, the GeneMark.hmm program that employed the typical and atypical models defined by GeneMark-Genesis was able to predict 683 new atypical genes with 176 of them confirmed by similarity search.

Algorithms↗

"A system biology" approach to bioinformatics and functional genomics in complex human diseases: arthritis.

Human and other annotated genome sequences have facilitated generation of vast amounts of correlative data, from human/animal genetics, normal and disease-affected tissues from complex diseases such as arthritis using gene/protein chips and SNP analysis. These data sets include genes/proteins whose functions are partially known at the cellular level or may be completely unknown (e.g. ESTs). Thus, genomic research has transformed molecular biology from "data poor" to "data rich" science, allowing further division into subpopulations of subcellular fractions, which are often given an "-omic" suffix. These disciplines have to converge at a systemic level to examine the structure and dynamics of cellular and organismal function. The challenge of characterizing ESTs linked to complex diseases is like interpreting sharp images on a blurred background and therefore requires a multidimensional screen for functional genomics ("functionomics") in tissues, mice and zebra fish model, which intertwines various approaches and readouts to study development and homeostasis of a system. In summary, the post-genomic era of functionomics will facilitate to narrow the bridge between correlative data and causative data by quaint hypothesis-driven research using a system approach integrating "intercoms" of interacting and interdependent disciplines forming a unified whole as described in this review for Arthritis.

Animals↗

A high-resolution map of transcription in the yeast genome.

There is abundant transcription from eukaryotic genomes unaccounted for by protein coding genes. A high-resolution genome-wide survey of transcription in a well annotated genome will help relate transcriptional complexity to function. By quantifying RNA expression on both strands of the complete genome of Saccharomyces cerevisiae using a high-density oligonucleotide tiling array, this study identifies the boundary, structure, and level of coding and noncoding transcripts. A total of 85% of the genome is expressed in rich media. Apart from expected transcripts, we found operon-like transcripts, transcripts from neighboring genes not separated by intergenic regions, and genes with complex transcriptional architecture where different parts of the same gene are expressed at different levels. We mapped the positions of 3' and 5' UTRs of coding genes and identified hundreds of RNA transcripts distinct from annotated genes. These nonannotated transcripts, on average, have lower sequence conservation and lower rates of deletion phenotype than protein coding genes. Many other transcripts overlap known genes in antisense orientation, and for these pairs global correlations were discovered: UTR lengths correlated with gene function, localization, and requirements for regulation; antisense transcripts overlapped 3' UTRs more than 5' UTRs; UTRs with overlapping antisense tended to be longer; and the presence of antisense associated with gene function. These findings may suggest a regulatory role of antisense transcription in S. cerevisiae. Moreover, the data show that even this well studied genome has transcriptional complexity far beyond current annotation.

5' Untranslated Regions↗

Systematic identification of pseudogenes through whole genome expression evidence profiling.

The identification of pseudogenes is an integral and significant part of the genome annotation because of their abundance and their impact on the experimental analysis of functional genes. Most of the computational annotation systems are not optimized for systematic pseudogene recognition, often annotating pseudogenes as functional genes, and users then propagate these errors to subsequent analyses and interpretations. In order to validate gene annotations and to identify pseudogenes that are potentially mis-annotated, we developed a novel approach based on whole genome profiling of existing transcript and protein sequences. This method has two important features: (i) equally detects both processed and non-processed pseudogenes and (ii) can identify transcribed pseudogenes. Applying this method to the human Ensembl gene predictions, we discovered that 2011 (9% of total) Ensembl genes in the categories of known and novel might be pseudogenes based on expression evidence. Of these, 1200 genes are found to have no existing evidence of transcription, and 811 genes are found with transcription evidence but contain significant translation disruption. Approximately 40% of the 2011 identified pseudogenes presented a multi-exon structure, representing non-processed pseudogenes. We have demonstrated the power of whole genome profiling of expression sequences to improve the accuracy of gene annotations.

Computational Biology↗

The complete set of tRNA species in Nanoarchaeum equitans.

The archaeal parasite Nanoarchaeum equitans was found to generate five tRNA species via a unique process requiring the assembly of seperate 5' and 3' tRNA halves [Randau, L., Munch, R., Hohn, M.J., Jahn, D. and Soll, D. (2005) Nanoarchaeum equitans creates functional tRNAs from separate genes for their 5'- and 3'-halves. Nature 433, 537-541]. Biochemical evidence was missing for one of the computationally-predicted, joined tRNAs designated as tRNA(Trp). Our RT-PCR and sequencing results identify this tRNA as tRNA(Lys) (CUU) joined at the alternative position between bases 30 and 31. We show that the intron-containing tRNA(Trp) was misidentified in the initial Nanoarchaeum equitans genome annotation [E. Waters et al. (2003) The genome of Nanoarchaeum equitans: insights into early archaeal evolution and derived parasitism. Proc. Natl. Acad. Sci. USA 100, 12984-12988]. Along with a previously unidentified joined tRNA(Gln) (UUG), Nanoarchaeum equitans exhibits 44 tRNAs and is enabled to read all 61 sense codons. Features unique to this set of tRNA molecules are discussed.

Base Sequence↗

Using proteomics to mine genome sequences.

We present a method for mining unannotated or annotated genome sequences with proteomic data to identify open reading frames. The region of a genome coding for a protein sequence is identified by using information from the analysis of proteins and peptides with MALDI-TOF mass spectrometry. The raw genome sequence or any unassembled contigs of an organism are theoretically cleaved into a number of equal sized but overlapping fragments, and these are then translated in all six frames into a series of virtual proteins. Each virtual protein is then subjected to a theoretical enzymatic digestion. Standard proteomic sample preparation methods are used to separate, array, and digest the proteins of interest to peptides. The masses of the resulting peptides are measured using mass spectrometry and compared to the theoretical peptide masses of the virtual proteins. The region of the genome responsible for coding for a particular protein can then be identified when there are a large number of hits between peptides from the protein and peptides from the virtual protein. The method makes no assumptions about the location of a protein in a particular gene sequence or the positions or types of start and stop codons. To illustrate this approach, all 773 proteins of Pseudomonas aeruginosa contained in SWISS-PROT were used to theoretically test the method and optimize parameters. Increasing the size of the virtual proteins results in an overall improvement in the ability to detect the coding region, at the cost of decreasing the sensitivity of the method for smaller proteins. Increasing the minimum number of matching peptides, lowering the mass error tolerance, or increasing the signal-to-noise ratio of the simulated mass spectrum, improves the ability to detect coding regions. The method is further demonstrated on experimental data from Mycobacterium tuberculosis and is also shown to work with eukaryotic organisms (e.g., Homo sapiens).

Amino Acid Sequence↗

EC_oligos: automated and whole-genome primer design for exons within one or between two genomes.

SUMMARY: EC_oligos designs oligonucleotides (oligos) from exons of annotated genomic sequence information. It can automatically and rapidly select oligos that are conserved between two sets of sequence data, and can pair up oligos for use as PCR primers. It can do this on a whole-genome scale and according to user-defined criteria. AVAILABILITY: The source code, executable program and user manual are available at ftp://ftp.ebi.ac.uk/pub/software/dos/EC_oligos/.

Algorithms↗

BuchneraBASE: a post-genomic resource for Buchnera sp. APS.

SUMMARY: BuchneraBASE is a bioinformatic research tool for the genome of the symbiotic bacterium Buchnera sp. APS that includes an improved genome annotation, comparative information about related insect symbiont genomes and a complete mapping of metabolic reactions to an Escherichia coli in silico model. The database is designed to accommodate genome-wide post-genomic datasets that are becoming available for this organism. AVAILABILITY: BuchneraBASE is available at http://www.buchnera.org/.

Buchnera↗

dictyBase, the model organism database for Dictyostelium discoideum.

dictyBase (http://dictybase.org) is the model organism database (MOD) for the social amoeba Dictyostelium discoideum. The unique biology and phylogenetic position of Dictyostelium offer a great opportunity to gain knowledge of processes not characterized in other organisms. The recent completion of the 34 MB genome sequence, together with the sizable scientific literature using Dictyostelium as a research organism, provided the necessary tools to create a well-annotated genome. dictyBase has leveraged software developed by the Saccharomyces Genome Database and the Generic Model Organism Database project. This has reduced the time required to develop a full-featured MOD and greatly facilitated our ability to focus on annotation and providing new functionality. We hope that manual curation of the Dictyostelium genome will facilitate the annotation of other genomes.

Animals↗

RRE: a tool for the extraction of non-coding regions surrounding annotated genes from genomic datasets.

UNLABELLED: RRE allows the extraction of non-coding regions surrounding a coding sequence [i.e. gene upstream region, 5'-untranslated region (5'-UTR), introns, 3'-UTR, downstream region] from annotated genomic datasets available at NCBI. AVAILABILITY: RRE parser and web-based interface are accessible at http://www.bioinformatica.unito.it/bioinformatics/rre/rre.html

Chromosome Mapping↗