Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “genome annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Reference based annotation with GeneMapper.

We introduce GeneMapper, a program for transferring annotations from a well annotated genome to other genomes. Drawing on high quality curated annotations, GeneMapper enables rapid and accurate annotation of newly sequenced genomes and is suitable for both finished and draft genomes. GeneMapper uses a profile based approach for mapping genes into multiple species, improving upon the standard pairwise approach. GeneMapper is freely available for academic use.

Algorithms↗

The Arabidopsis genome sequence as a tool for genome analysis in Brassicaceae. A comparison of the Arabidopsis and Capsella rubella genomes.

The annotated Arabidopsis genome sequence was exploited as a tool for carrying out comparative analyses of the Arabidopsis and Capsella rubella genomes. Comparison of a set of random, short C. rubella sequences with the corresponding sequences in Arabidopsis revealed that aligned protein-coding exon sequences differ from aligned intron or intergenic sequences in respect to the degree of sequence identity and the frequency of small insertions/deletions. Molecular-mapped markers and expressed sequence tags derived from Arabidopsis were used for genetic mapping in a population derived from an interspecific cross between Capsella grandiflora and C. rubella. The resulting eight Capsella linkage groups were compared to the sequence maps of the five Arabidopsis chromosomes. Fourteen colinear segments spanning approximately 85% of the Arabidopsis chromosome sequence maps and 92% of the Capsella genetic linkage map were detected. Several fusions and fissions of chromosomal segments as well as large inversions account for the observed arrangement of the 14 colinear blocks in the analyzed genomes. In addition, evidence for small-scale deviations from genome colinearity was found. Colinearity between the Arabidopsis and Capsella genomes is more pronounced than has been previously reported for comparisons between Arabidopsis and different Brassica species.

Arabidopsis↗

Large-scale mutational analysis for the annotation of the mouse genome.

After sequencing the human and mouse genomes, the annotation of these sequences with biological functions is an important challenge in genomic research. A major tool to analyse gene function on the organismal level is the analysis of mutant phenotypes. Because of its genetic and physiological similarity to man, the mouse has become the model organism of choice for the study of genetic diseases. In addition, there is at the moment no other vertebrate for which versatile techniques to manipulate the genome are as well developed. Several mouse mutagenesis projects have provided the proof-of-principle that a systematic and comprehensive mutagenesis of every gene in the mammalian genome will be feasible. An exhaustive functional annotation of the mammalian genome can only be achieved in a combination of phenotype- and gene-driven approaches in large- and small-scale academic and private projects. Major challenges will be to develop standardised phenotyping protocols for the clinical and pathological characterisation of mouse mutants, the improvement of mutation detection methods and the dissemination of resources and data. Beyond gene annotation, it will be necessary to understand how gene functions are integrated into the complex network of regulatory interactions in the cell.

Animals↗

Chromosome-level genome assembly of Sinocyclocheilus jii based on PacBio HiFi and Hi-C sequencing.

Sinocyclocheilus jii, a cavefish species endemic to China, belongs to the genus Sinocyclocheilus within the family Cyprinidae. Species within this genus exhibit significant morphological differentiation, making it not only the most species-rich genus within Cyprinidae in China but also the most diverse group of cavefishes worldwide. However, the limited availability of genomic resources has limited investigations into the genetic basis of trait variations, phylogenetic relationships, and adaptive evolution in this genus. In this study, we assembled a chromosome-level reference genome for S. jii by integrating PacBio HiFi long reads, Illumina short reads, and Hi-C sequencing data. Flow cytometry was used to estimate the genome size prior to assembly, providing a key step in technical validation. The final genome assembly spans 1.75 Gb with a contig N50 of 35.0 Mb. Using Hi-C sequencing data, the assembled scaffolds were successfully anchored to 50 chromosomes. The completeness of the chromosome-level assembly was estimated at 98.9% by BUSCO analysis. Genome annotation identified 855.5 Mb of repetitive sequences and predicted a total of 52,867 protein-coding genes, of which 51,932 genes were functionally annotated. This study presents a high-quality chromosome-level genome assembly and annotation of S. jii, providing a fundamental genomic resource for future phylogenetic and evolutionary studies.

Animals↗

Chromosome-level genome assembly and annotation of Pterygoplichthys pardalis.

Suckermouth catfishes, with their evolved powerful features, have become notorious invasive species, causing significant damage to aquatic ecosystems. However, the lack of high-quality genomes severely restricts research on this group within the field. In this study, we de novo assembled the chromosome-level genome assembly of Pterygoplichthys pardalis using multiple platforms of sequencing data, including Illumina short reads, Nanopore long reads, and Hi-C sequencing reads, resulting in a 1.51 Gb genome assembly. Multiple evaluations, including read mapping ratio (98.52%), transcript mapping ratio (99.61%), conserved BUSCO gene set (98.8%), and N50 score (49.47 Mb), indicated the high continuity and accuracy of the genome assembly we generated. Genome annotation found that 0.97 Gb of genome sequences are repetitive sequences, accounting for 64.47% of the genome assembly. Further, 23,859 protein-coding genes were successfully predicted, 92.92% of which could be annotated in functional databases. This high-quality genome assembly of P. pardalis provides a valuable resource for understanding the genetic underpinnings of P. pardalis's invasive success and offers critical data for future fisheries research and management.

Animals↗

In silico insight into two rice chromosomal regions associated with submergence tolerance and resistance to bacterial leaf blight and gall midge.

Plants respond to both biotic and abiotic stresses through a common signaling system to provide defense and protection against many adverse environments. Many genes/QTLs governing resistance to both biotic and abiotic stresses have been studied and mapped in rice. Sub1, a major QTL for submergence tolerance is collocated with a gene Gm1 for gall midge resistance on chromosome 9 (Region 1). Likewise a bigger region on chromosome 5 (Region 2) has a minor QTL for submergence tolerance collocated with genes for bacterial blight resistance. Utilizing the rice sequence and annotation data (TIGR) and rice genome annotation project database (RAP-DB), we wanted to know the kinds of genes underlying these two chromosomal regions where genes/QTL governing tolerance to both biotic and abiotic stresses are collocated. We also analyzed the pattern of distribution of these genes across the BAC/PAC clones spanning the region so that candidate genes can be short listed for a functional analysis. Genes known to have a role in submergence tolerance were present in both the regions. Region 1, had a unique transcription factor like trithorax protein, which is a positional candidate gene for submergence tolerance. Pyruvate decarboxylase (PDC) gene for alcohol fermentation and cation transporting ATPase c-terminal domain are likely candidates for submergence QTL in Region 2. Genes such as SKP1 and elicitor induced cytochrome p450 associated with tissue necrosis and insect resistance were found in region 1. Multiple copies of ORFs for signal transduction proteins, transcription factors, genes for systemic acquired resistance, Ubiquitin proteins and pathogen elicitor identification and degrading proteins were located as a cluster in Region 2, where bacterial blight resistance genes mapped. Validation of the data obtained from TIGR with other databases (RAP and KOME) confirmed our findings. The functional role of some of the significant candidate genes needs to be established. Allele/gene specific markers can then be designed for use in MAS thus enhancing durable tolerance/resistance faster.

Adaptation, Physiological↗

Remodelling of the homeobox gene complement in the tunicate Oikopleura dioica.

Homeodomain transcription factors are involved in many developmental processes and have been intensely studied in a few model organisms, such as mouse, Drosophila and Caenorhabditis elegans. Homeobox genes fall into 10 classes (ANTP, PRD, POU, LIM, TALE, SIX, Cut, ZFH, HNF1, Prox) and 89 different families/groups, all of which are present in vertebrates. Additional groups may be uncovered by further genome annotation, particularly of complex vertebrate genomes. Eight of these groups have been found only in vertebrates, but not in the genome of the tunicate Ciona intestinalis. The other 81 groups of homeobox gene that have been detected in vertebrates so far probably appeared during the early evolution of bilaterians or earlier, as they are also present outside the chordates. How the homeobox genes evolved during and after the main radiation of the bilaterians remains poorly understood, as only a few animal genomes have been sequenced completely. However, drastic changes have occurred at least in the lineage of C. elegans , such as loss of several Hox genes and Hox cluster fragmentation . Here we report considerable alterations of the homeobox gene complement in the tunicate lineage.

Animals↗

Systems properties of the Haemophilus influenzae Rd metabolic genotype.

Haemophilus influenzae Rd was the first free-living organism for which the complete genomic sequence was established. The annotated sequence and known biochemical information was used to define the H. influenzae Rd metabolic genotype. This genotype contains 488 metabolic reactions operating on 343 metabolites. The stoichiometric matrix was used to determine the systems characteristics of the metabolic genotype and to assess the metabolic capabilities of H. influenzae. The need to balance cofactor and biosynthetic precursor production during growth on mixed substrates led to the definition of six different optimal metabolic phenotypes arising from the same metabolic genotype, each with different constraining features. The effects of variations in the metabolic genotype were also studied, and it was shown that the H. influenzae Rd metabolic genotype contains redundant functions under defined conditions. We thus show that the synthesis of in silico metabolic genotypes from annotated genome sequences is possible and that systems analysis methods are available that can be used to analyze and interpret phenotypic behavior of such genotypes.

Cell Division↗

PLOTREP: a web tool for defragmentation and visual analysis of dispersed genomic repeats.

Identification of dispersed or interspersed repeats, most of which are derived from transposons, retrotransposons or retrovirus-like elements, is an important step in genome annotation. Software tools that compare genomic sequences with precompiled repeat reference libraries using sensitive similarity-based methods provide reliable means of finding the positions of fragments homologous to known repeats. However, their output is often incomplete and fragmented owing to the mutations (nucleotide substitutions, deletions or insertions) that can result in considerable divergence from the reference sequence. Merging these fragments to identify the whole region that represents an ancient copy of a mobile element is challenging, particularly if the element is large and suffered multiple deletions or insertions. Here we report PLOTREP, a tool designed to post-process results obtained by sequence similarity search and merge fragments belonging to the same copy of a repeat. The software allows rapid visual inspection of the results using a dot-plot like graphical output. The web implementation of PLOTREP is available at http://bioinformatics.abc.hu/PLOTREP/.

Computer Graphics↗

BeetleBase: the model organism database for Tribolium castaneum.

BeetleBase (http://www.bioinformatics.ksu.edu/BeetleBase/) is an integrated resource for the Tribolium research community. The red flour beetle (Tribolium castaneum) is an important model organism for genetics, developmental biology, toxicology and comparative genomics, the genome of which has recently been sequenced. BeetleBase is constructed to integrate the genomic sequence data with information about genes, mutants, genetic markers, expressed sequence tags and publications. BeetleBase uses the Chado data model and software components developed by the Generic Model Organism Database (GMOD) project. This strategy not only reduces the time required to develop the database query tools but also makes the data structure of BeetleBase compatible with that of other model organism databases. BeetleBase will be useful to the Tribolium research community for genome annotation as well as comparative genomics.

Animals↗

HmtDB, a human mitochondrial genomic resource based on variability studies supporting population genetics and biomedical research.

BACKGROUND: Population genetics studies based on the analysis of mtDNA and mitochondrial disease studies have produced a huge quantity of sequence data and related information. These data are at present worldwide distributed in differently organised databases and web sites not well integrated among them. Moreover it is not generally possible for the user to submit and contemporarily analyse its own data comparing them with the content of a given database, both for population genetics and mitochondrial disease data. RESULTS: HmtDB is a well-integrated web-based human mitochondrial bioinformatic resource aimed at supporting population genetics and mitochondrial disease studies, thanks to a new approach based on site-specific nucleotide and aminoacid variability estimation. HmtDB consists of a database of Human Mitochondrial Genomes, annotated with population data, and a set of bioinformatic tools, able to produce site-specific variability data and to automatically characterize newly sequenced human mitochondrial genomes. A query system for the retrieval of genomes and a web submission tool for the annotation of new genomes have been designed and will soon be implemented. The first release contains 1255 fully annotated human mitochondrial genomes. Nucleotide site-specific variability data and multialigned genomes can be downloaded. Intra-human and inter-species aminoacid variability data estimated on the 13 coding for proteins genes of the 1255 human genomes and 60 mammalian species are also available. HmtDB is freely available, upon registration, at http://www.hmdb.uniba.it. CONCLUSION: The HmtDB project will contribute towards completing and/or refining haplogroup classification and revealing the real pathogenic potential of mitochondrial mutations, on the basis of variability estimation.

Computational Biology↗

Streamlining large-scale genomic data management: Insights from the UK Biobank whole-genome sequencing data.

Biobank-scale whole-genome sequencing (WGS) studies are increasingly pivotal in unraveling the genetic bases of diverse health outcomes. However, managing and analyzing these datasets' sheer volume and complexity presents significant challenges. We highlight the annotated genomic data structure (aGDS) format, substantially reducing the WGS data file size while enabling seamless integration of genomic and functional information for comprehensive WGS analyses. The aGDS format yielded 23 chromosome-specific files for the UK Biobank 500k WGS dataset, occupying only 1.10 tebibytes of storage. We develop the vcf2agds toolkit that streamlines the conversion of WGS data from VCF to aGDS format. Additionally, the STAARpipeline equipped with the aGDS files enabled scalable, comprehensive, and functionally informed WGS analysis, facilitating the detection of common and rare coding and noncoding phenotype-genotype associations. Overall, the vcf2agds toolkit and STAARpipeline provide a streamlined solution that facilitates efficient data management and analysis of biobank-scale WGS data across hundreds of thousands of samples.

Humans↗

Computational analysis of Plasmodium falciparum metabolism: organizing genomic information to facilitate drug discovery.

Identification of novel targets for the development of more effective antimalarial drugs and vaccines is a primary goal of the Plasmodium genome project. However, deciding which gene products are ideal drug/vaccine targets remains a difficult task. Currently, a systematic disruption of every single gene in Plasmodium is technically challenging. Hence, we have developed a computational approach to prioritize potential targets. A pathway/genome database (PGDB) integrates pathway information with information about the complete genome of an organism. We have constructed PlasmoCyc, a PGDB for Plasmodium falciparum 3D7, using its annotated genomic sequence. In addition to the annotations provided in the genome database, we add 956 additional annotations to proteins annotated as "hypothetical" using the GeneQuiz annotation system. We apply a novel computational algorithm to PlasmoCyc to identify 216 "chokepoint enzymes." All three clinically validated drug targets are chokepoint enzymes. A total of 87.5% of proposed drug targets with biological evidence in the literature are chokepoint reactions. Therefore, identifying chokepoint enzymes represents one systematic way to identify potential metabolic drug targets.

Algorithms↗

The Genomic Threading Database: a comprehensive resource for structural annotations of the genomes from key organisms.

Currently, the Genomic Threading Database (GTD) contains structural assignments for the proteins encoded within the genomes of nine eukaryotes and 101 prokaryotes. Structural annotations are carried out using a modified version of GenTHREADER, a reliable fold recognition method. The Gen THREADER annotation jobs are distributed across multiple clusters of processors using grid technology and the predictions are deposited in a relational database accessible via a web interface at http://bioinf.cs.ucl.ac.uk/GTD. Using this system, up to 84% of proteins encoded within a genome can be confidently assigned to known folds with 72% of the residues aligned. On average in the GTD, 64% of proteins encoded within a genome are confidently assigned to known folds and 58% of the residues are aligned to structures.

Animals↗

GenomeRNAi: a database for cell-based RNAi phenotypes.

RNA interference (RNAi) has emerged as a powerful tool to generate loss-of-function phenotypes in a variety of organisms. Combined with the sequence information of almost completely annotated genomes, RNAi technologies have opened new avenues to conduct systematic genetic screens for every annotated gene in the genome. As increasing large datasets of RNAi-induced phenotypes become available, an important challenge remains the systematic integration and annotation of functional information. Genome-wide RNAi screens have been performed both in Caenorhabditis elegans and Drosophila for a variety of phenotypes and several RNAi libraries have become available to assess phenotypes for almost every gene in the genome. These screens were performed using different types of assays from visible phenotypes to focused transcriptional readouts and provide a rich data source for functional annotation across different species. The GenomeRNAi database provides access to published RNAi phenotypes obtained from cell-based screens and maps them to their genomic locus, including possible non-specific regions. The database also gives access to sequence information of RNAi probes used in various screens. It can be searched by phenotype, by gene, by RNAi probe or by sequence and is accessible at http://rnai.dkfz.de.

Animals↗

EXProt: a database for proteins with an experimentally verified function.

EXProt is a non-redundant protein database containing a selection of entries from genome annotation projects and public databases, aimed at including only proteins with an experimentally verified function. In EXProt release 2.0 we have collected entries from the Pseudomonas aeruginosa community annotation project (PseudoCAP), the Escherichia coli genome and proteome database (GenProtEC) and the translated coding sequences from the Prokaryotes division of EMBL nucleotide sequence database, which are described as having an experimentally verified function. Each entry in EXProt has a unique ID number and contains information about the species, amino acid sequence, functional annotation and, in most cases, links to references in MEDLINE/PubMed and to the entry in the original database. EXProt is indexed in SRS at CMBI (http://www.cmbi.kun.nl/srs/) and can be searched with BLAST and FASTA through the EXProt web page (http://www.cmbi.kun.nl/EXProt/).

Animals↗

Defining genes in the genome of the hyperthermophilic archaeon Pyrococcus furiosus: implications for all microbial genomes.

The original genome annotation of the hyperthermophilic archaeon Pyrococcus furiosus contained 2,065 open reading frames (ORFs). The genome was subsequently automatically annotated in two public databases by the Institute for Genomic Research (TIGR) and the National Center for Biotechnology Information (NCBI). Remarkably, more than 500 of the originally annotated ORFs differ in size in the two databases, many very significantly. For example, more than 170 of the predicted proteins differ at their N termini by more than 25 amino acids. Similar discrepancies were observed in the TIGR and NCBI databases with the other archaeal and bacterial genomes examined. In addition, the two databases contain 60 (NCBI) and 221 (TIGR) ORFs not present in the original annotation of P. furiosus. In the present study we have experimentally assessed the validity of 88 previously unannotated ORFs. Transcriptional analyses showed that 11 of 61 ORFs examined were expressed in P. furiosus when grown at either 95 or 72 degrees C. In addition, 7 of 54 ORFs examined yielded heat-stable recombinant proteins when they were expressed in Escherichia coli, although only one of the seven ORFs was expressed in P. furiosus under the growth conditions tested. It is concluded that the P. furiosus genome contains at least 17 ORFs not previously recognized in the original annotation. This study serves to highlight the discrepancies in the public databases and the problems of accurately defining the number and sizes of ORFs within any microbial genome.

Archaeal Proteins↗

GeneFizz: A web tool to compare genetic (coding/non-coding) and physical (helix/coil) segmentations of DNA sequences. Gene discovery and evolutionary perspectives.

The GeneFizz (http://pbga.pasteur.fr/GeneFizz) web tool permits the direct comparison between two types of segmentations for DNA sequences (possibly annotated): the coding/non-coding segmentation associated with genomic annotations (simple genes or exons in split genes) and the physics-based structural segmentation between helix and coil domains (as provided by the classical helix-coil model). There appears to be a varying degree of coincidence for different genomes between the two types of segmentations, from almost perfect to non-relevant. Following these two extremes, GeneFizz can be used for two purposes: ab initio physics-based identification of new genes (as recently shown for Plasmodium falciparum) or the exploration of possible evolutionary signals revealed by the discrepancies observed between the two types of information.

Algorithms↗