Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “genome annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

OmicBrowse: a browser of multidimensional omics annotations.

UNLABELLED: OmicBrowse is a browser to explore multiple datasets coordinated in the multidimensional omic space integrating omics knowledge ranging from genomes to phenomes and connecting evolutional correspondences among multiple species. OmicBrowse integrates multiple data servers into a single omic space through secure peer-to-peer server communications, so that a user can easily obtain an integrated view of distributed data servers, e.g. an integrated view of numerous whole-genome tiling-array data retrieved from a user's in-house private-data server, along with various genomic annotations from public internet servers. OmicBrowse is especially appropriate for positional-cloning purposes. It displays both genetic maps and genomic annotations within wide chromosomal intervals and assists a user to select candidate genes by filtering their annotations or associated documents against user-specified keywords or ontology terms. We also show that an omic-space chart effectively represents schemes for integrating multiple datasets of multiple species. AVAILABILITY: OmicBrowse is developed by the Genome-Phenome Superbrain Project and is released as free open-source software under the GNU General Public License at http://omicspace.riken.jp.

Chromosome Mapping↗

The presence of GC-AG introns in Neurospora crassa and other euascomycetes determined from analyses of complete genomes: implications for automated gene prediction.

A combination of experimental and computational approaches was employed to identify introns with noncanonical GC-AG splice sites (GC-AG introns) within euascomycete genomes. Evaluation of 2335 cDNA-confirmed introns from Neurospora crassa revealed 27 such introns (1.2%). A similar frequency (1.0%) of GC-AG introns was identified in Fusarium graminearum, in which 3 of 292 cDNA-confirmed introns contained GC-AG splice sites. Computational analyses of the N. crassa genome using a GC-AG intron consensus sequence identified an additional 20 probable GC-AG introns in this fungus. For 8 of the 47 GC-AG introns identified in N. crassa a GC donor site is also present in a homolog from Magnaporthe grisea, F. graminearum, or Aspergillus nidulans. In most cases, however, homologs in these fungi contain a GT-AG intron or no intron at the corresponding position. These findings have important implications for fungal genome annotation, as the automated annotations of euascomycete genomes incorrectly identified intron boundaries for all of the confirmed and probable GC-AG introns reported here.

Alternative Splicing↗

Proteome analysis of an aerobic hyperthermophilic crenarchaeon, Aeropyrum pernix K1.

We analyzed the proteome of a crenararchaeon, Aeropyrum pernix K1, by using the following four methods: (i) two-dimensional PAGE followed by MALDI-TOF MS, (ii) one-dimensional SDS-PAGE in combination with two-dimensional LC-MS/MS, (iii) multidimensional LC-MS/MS, and (iv) two-dimensional PAGE followed by amino-terminal amino acid sequencing. These methods were found to be complementary to each other, and biases in the data obtained in one method could largely be compensated by the data obtained in the other methods. Consequently a total of 704 proteins were successfully identified, 134 of which were unique to A. pernix K1, and 19 were not described previously in the genomic annotation. We found that the original annotation of the genomic data of this archaeon was not adequate in particular with respect to proteins of 10-20 kDa in size, many of which were described as hypothetical. Furthermore the amino-terminal amino acid sequence analysis indicated that surprisingly the translation of 52% of their genes starts with TTG in contrast to ATG (28%) and GTG (20%). Thus, A. pernix K1 is the first example of an organism in which TTG is the most predominant translational initiation codon.

Aerobiosis↗

Target Explorer: An automated tool for the identification of new target genes for a specified set of transcription factors.

With the increasing number of eukaryotic genomes available, high-throughput automated tools for identification of regulatory DNA sequences are becoming increasingly feasible. Several computational approaches for the prediction of regulatory elements were recently developed. Here we combine the prediction of clusters of binding sites for transcription factors with context information taken from genome annotations. Target Explorer automates the entire process from the creation of a customized library of binding sites for known transcription factors through the prediction and annotation of putative target genes that are potentially regulated by these factors. It was specifically designed for the well-annotated Drosophila melanogaster genome, but most options can be used for sequences from other genomes as well. Target Explorer is available at http://trantor.bioc.columbia.edu/Target_Explorer/

Animals↗

Whole blood transcriptome profile identifies motor neurone disease RNA biomarker signatures.

Blood-based biomarkers for motor neuron disease are needed for better diagnosis, progression prediction, and clinical trial monitoring. We used whole blood-derived total RNA and performed whole transcriptome analysis to compare the gene expression profiles in (motor neurone disease) MND patients to the control subjects. We compared 42 MND patients to 42 aged and sex-matched healthy controls and described the whole transcriptome profile characteristic for MND. In addition to the formal differential analysis, we performed functional annotation of the genomics data and identified the molecular pathways that are differentially regulated in MND patients. We identified 12,972 genes differentially expressed in the blood of MND patients compared to age and sex-matched controls. Functional genomic annotation identified activation of the pathways related to neurodegeneration, RNA transcription, RNA splicing and extracellular matrix reorganisation. Blood-based whole transcriptomic analysis can reliably differentiate MND patients from controls and can provide useful information for the clinical management of the disease and clinical trials.

Humans↗

OrthoMCL: identification of ortholog groups for eukaryotic genomes.

The identification of orthologous groups is useful for genome annotation, studies on gene/protein evolution, comparative genomics, and the identification of taxonomically restricted sequences. Methods successfully exploited for prokaryotic genome analysis have proved difficult to apply to eukaryotes, however, as larger genomes may contain multiple paralogous genes, and sequence information is often incomplete. OrthoMCL provides a scalable method for constructing orthologous groups across multiple eukaryotic taxa, using a Markov Cluster algorithm to group (putative) orthologs and paralogs. This method performs similarly to the INPARANOID algorithm when applied to two genomes, but can be extended to cluster orthologs from multiple species. OrthoMCL clusters are coherent with groups identified by EGO, but improved recognition of "recent" paralogs permits overlapping EGO groups representing the same gene to be merged. Comparison with previously assigned EC annotations suggests a high degree of reliability, implying utility for automated eukaryotic genome annotation. OrthoMCL has been applied to the proteome data set from seven publicly available genomes (human, fly, worm, yeast, Arabidopsis, the malaria parasite Plasmodium falciparum, and Escherichia coli). A Web interface allows queries based on individual genes or user-defined phylogenetic patterns (http://www.cbil.upenn.edu/gene-family). Analysis of clusters incorporating P. falciparum genes identifies numerous enzymes that were incompletely annotated in first-pass annotation of the parasite genome.

Animals↗

The complete and annotated mitochondrial genome of Hemileia vastatrix Race I, causal agent of coffee leaf rust.

Hemileia vastatrix is the fungal pathogen responsible for coffee leaf rust (CLR), the most economically important disease of Coffea arabica worldwide. Recently, the nuclear genome of this fungus was completely deciphered. However, the mitochondrial genome of H. vastatrix has remained undercharacterized. Here, we present the complete, circularized mitochondrial genome of H. vastatrix Race I (isolate HvRI), assembled using a hybrid approach combining PacBio HiFi long reads and BGIseq short reads. The genome is 173,525 bp in length with a GC content of 33.1% and encodes 41 functional genes, including 15 protein-coding genes, 2 rRNAs, and 24 tRNAs. The assembly reveals significant structural complexity, driven by intron expansion in the cox1 and cob genes. Notably, the atp8 gene contains a group II intron, rare for this locus, whose internal open reading frame displays evidence of pseudogenization via internal stop codons.. We also characterized a putative replication initiation zone (~1.2 kb) defined by a poly-G homopolymer and conserved regulatory motifs. The mitogenome of the HvRI isolate does not contain cob mutations that lead to amino acid substitutions G143A and F129L associated with the quinone outside inhibitor (QoI) fungicide resistance. This high-quality mitogenome is an important resource for comparative mitogenomics, population diversity studies, and the molecular surveillance of QoI fungicide resistance.

Genome, Mitochondrial↗

The Drosophila melanogaster genome sequencing and annotation projects: a status report.

The sequence and genome annotations of Drosophila melanogaster were initially published in late 1999 and early 2000. Since then, the Berkeley Drosophila Genome Project (BDGP) and FlyBase have improved the quality of the sequence and reviewed the annotations by hand, respectively, to produce an account of the fruit fly genome that is of the highest quality. This review discusses the main features of this process, both from the point of view of the biology revealed in the end result and in the development of software that has been central to this genome sequencing and annotation project.

Animals↗

Genome-tools: a flexible package for genome sequence analysis.

Genome-tools is a Perl module, a set of programs, and a user interface that facilitates access to genome sequence information. The package is flexible, extensible, and designed to be accessible and useful to both nonprogrammers and programmers. Any relatively well-annotated genome available with standard GenBank genome files may be used with genome-tools. A simple Web-based front end permits searching any available genome with an intuitive interface. Flexible design choices also make it simple to handle revised versions of genome annotation files as they change. In addition, programmers can develop cross-genomic tools and analyses with minimal additional overhead by combining genome-tools modules with newly written modules. Genome-tools runs on any computer platform for which Perl is available, including Unix, Microsoft Windows, and Mac OS. By simplifying the access to large amounts of genomic data, genome-tools may be especially useful for molecular biologists looking at newly sequenced genomes, for which few informatics tools are available. The genome-tools Web interface is accessible at http://genome-tools.sourceforge.net, and the source code is available at http://sourceforge.net/projects/genome-tools.

Base Sequence↗

The distributed annotation system.

BACKGROUND: Currently, most genome annotation is curated by centralized groups with limited resources. Efforts to share annotations transparently among multiple groups have not yet been satisfactory. RESULTS: Here we introduce a concept called the Distributed Annotation System (DAS). DAS allows sequence annotations to be decentralized among multiple third-party annotators and integrated on an as-needed basis by client-side software. The communication between client and servers in DAS is defined by the DAS XML specification. Annotations are displayed in layers, one per server. Any client or server adhering to the DAS XML specification can participate in the system; we describe a simple prototype client and server example. CONCLUSIONS: The DAS specification is being used experimentally by Ensembl, WormBase, and the Berkeley Drosophila Genome Project. Continued success will depend on the readiness of the research community to adopt DAS and provide annotations. All components are freely available from the project website http://www.biodas.org/.

Base Sequence↗

The MitoDrome database annotates and compares the OXPHOS nuclear genes of Drosophila melanogaster, Drosophila pseudoobscura and Anopheles gambiae.

The oxidative phosphorylation (OXPHOS) is the primary energy-producing process of all aerobic organisms and the only cellular function under the dual control of both the mitochondrial and the nuclear genomes. Functional characterization and evolutionary study of the OXPHOS system is of great importance for the understanding of many as yet unclear aspects of nucleus-mitochondrion genomic co-evolution and co-regulation gene networks. The MitoDrome database is a web-based database which provides genomic annotations about nuclear genes of Drosophila melanogaster encoding for mitochondrial proteins. Recently, MitoDrome has included a new section annotating genomic information about OXPHOS genes in Drosophila pseudoobscura and Anopheles gambiae and their comparative analysis with their Drosophila melanogaster and human counterparts. The introduction of this new comparative annotation section into MitoDrome is expected to be a useful resource for both functional and structural genomics related to the OXPHOS system.

Animals↗

A genome annotation-driven approach to cloning the human ORFeome.

We have developed a systematic approach to generating cDNA clones containing full-length open reading frames (ORFs), exploiting knowledge of gene structure from genomic sequence. Each ORF was amplified by PCR from a pool of primary cDNAs, cloned and confirmed by sequencing. We obtained clones representing 70% of genes on human chromosome 22, whereas searching available cDNA clone collections found at best 48% from a single collection and 60% for all collections combined.

Chromosomes, Human, Pair 22↗

U12DB: a database of orthologous U12-type spliceosomal introns.

U12-type introns are spliced by the U12-dependent spliceosome and are present in the genomes of many higher eukaryotic lineages including plants, chordates and some invertebrates. However, due to their relatively recent discovery and a systematic bias against recognition of non-canonical splice sites in general, the introns defined by U12-type splice sites are under-represented in genome annotations. Such under-representation compounds the already difficult problem of determining gene structures. It also impedes attempts to study these introns genome-wide or phylum-wide. The resource described here, the U12 Intron Database (U12DB), aims to catalog the U12-type introns of completely sequenced eukaryotic genomes in a framework that groups orthologous introns with each other. This will aid further investigations into the evolution and mechanism of U12-dependent splicing as well as assist ongoing genome annotation efforts. Public access to the U12DB is available at http://genome.imim.es/cgi-bin/u12db/u12db.cgi.

Animals↗

Genome-wide annotation of human multi-nucleotide variants reveals widespread functional differences from single nucleotide variants.

Multi-nucleotide variants (MNVs) represent a crucial yet underexplored category of genetic variation. Despite previous studies highlighting the prevalence and potential biological impact of MNVs in populations, comprehensive identification and detailed functional annotation of MNVs remain challenging. Here, we develop MNVAnno, a toolbox for rapid identification and annotation of complex MNVs, and utilize it to identify 3,984,258 MNVs from 700,134 human samples, expanding the human MNV list to 8,199,654. Our analysis reveals that MNVs can not only lead to distinct amino acid changes from their constituent single-nucleotide variants, but also significantly impact the function of non-coding regions. Furthermore, through genome-wide association studies, we identify some MNVs associated with multiple cancers, and establish the Human MNV Database to facilitate MNV research. Our study emphasizes the importance of MNV annotation, broadens the human MNV landscape, and opens avenues for exploring genetic variation in phenotypes and diseases.

Humans↗

Bioinformatics issues for automating the annotation of genomic sequences.

The rapid explosion in the amount of biological data being generated worldwide is surpassing efforts to manage analysis of the data. As part of an ongoing project to automate and manage bioinformatics analysis, the authors have designed and implemented a simple automated annotation system, which is described in this paper. The system is applied to existing GenBank/DDBJ/EMBL entries and compared with existing annotations to illustrate not only potential errors but also that they are generally not up-to-date, as a result of new versions of analysis tools and updates of genomic repositories. We highlight the important Bioinformatics issues of storage and management of information to ensure data and results are kept up-to-date in light of new information becoming available. Surprisingly, from just four database entries, a significant number of new features were found. We describe the results as well as identify important issues that need to be addressed in order to automate the re-analysis/re-annotation of genomic sequences within a reasonable timeframe.

Computational Biology↗