Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “genome annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Checks and balancers: balancer chromosomes to facilitate genome annotation.

Phenotype-driven mutagenesis screens are used to discover gene function in model organisms. Mutations that are induced by chemical mutagens can occur anywhere in the genome. However, the use of a balancer chromosome (where a phenotypically marked segment of a chromosome is inverted) in a mutagenesis screen enables mutations to be mapped in a defined region of the genome and maintained stably in a heterozygous state. Mouse balancer chromosomes can be engineered using Cre-loxP technology in selected regions of the genome. Balancer mutagenesis screens will provide a systematic functional analysis of the genes on mouse chromosomes, and consequently, will facilitate a functional annotation of the mammalian genome sequence.

Animals↗

Annotating eukaryote genomes.

The Genome Annotation Assessment Project tested current methods of gene identification, including a critical assessment of the accuracy of different methods. Two new databases have provided new resources for gene annotation: these are the InterPro database of protein domains and motifs, and the Gene Ontology database for terms that describe the molecular functions and biological roles of gene products. Efforts in genome annotation are most often based upon advances in computer systems that are specifically designed to deal with the tremendous amounts of data being generated by current sequencing projects. These efforts in analysis are being linked to new ways of visualizing computationally annotated genomes.

Animals↗

Representation and processing of complex DNA spatial architecture and its annotated genomic content.

This paper presents a new general approach for the spatial representation and visualization of DNA molecule and its annotated information. This approach is based on a biological 3D model that predicts the complex spatial trajectory of huge naked DNA. With such modeling, a global vision of the sequence is possible, which is different and complementary to other representations as textual, linguistics or syntactic ones. The DNA is well known as a three-dimensional structure. Whereas, the spatial information plays a great part during its evolution and its interaction with the other biological elements This work will motivate investigations in order to launch new bioinformatics studies for the analysis of the spatial architecture of the genome. Besides, in order to obtain a friendly interactive visualization, a powerful graphic modeling is proposed including DNA complex trajectory management and its annotated-based content structuring. The paper describes spatial architecture modeling, with consideration of both biological and computational constraints. This work is implemented through a powerful graphic software tool, named ADN-Viewer. Several examples of visualization are shown for various organisms and biological elements.

DNA↗

Genomic annotation and expression analysis of the zebrafish Rho small GTPase family during development and bacterial infection.

The zebrafish genomic sequence database was analyzed for the presence of genes encoding members of the Rho small GTPases. The analysis shows the presence of 32 zebrafish Rho genes representing one or more homologs of the human RHOA, RND3, RHOF, RHOG, RHOH, RHOJ, RHOU, RHOV, CDC42, RAC1, RAC2, RAC3, RND1, RHOBTB1, RHOBTB2, RHOBTB3, and RHOT1 genes. By expression analysis using reverse transcriptase-PCR we show that at least 20 of the predicted zebrafish small GTPase genes are expressed in the adult stage. Interestingly, only 5 of these were found to be expressed at early embryonic stages, including rhoab, rhoad, cdc42a, cdc42c, and rac1a. We observed a strong upregulation of zebrafish rhogb expression after Mycobacterium marinum infection of adult fish. This complete annotation study provides a firm basis for the use of zebrafish as a model for analysis of Rho GTPase function in vertebrate development and the innate immune system.

Amino Acid Sequence↗

Iterative gene prediction and pseudogene removal improves genome annotation.

Correct gene prediction is impaired by the presence of processed pseudogenes: nonfunctional, intronless copies of real genes found elsewhere in the genome. Gene prediction programs frequently mistake processed pseudogenes for real genes or exons, leading to biologically irrelevant gene predictions. While methods exist to identify processed pseudogenes in genomes, no attempt has been made to integrate pseudogene removal with gene prediction, or even to provide a freestanding tool that identifies such erroneous gene predictions. We have created PPFINDER (for Processed Pseudogene finder), a program that integrates several methods of processed pseudogene finding in mammalian gene annotations. We used PPFINDER to remove pseudogenes from N-SCAN gene predictions, and show that gene prediction improves substantially when gene prediction and pseudogene masking are interleaved. In addition, we used PPFINDER with gene predictions as a parent database, eliminating the need for libraries of known genes. This allows us to run the gene prediction/PPFINDER procedure on newly sequenced genomes for which few genes are known.

Animals↗

Genome annotation past, present, and future: how to define an ORF at each locus.

Driven by competition, automation, and technology, the genomics community has far exceeded its ambition to sequence the human genome by 2005. By analyzing mammalian genomes, we have shed light on the history of our DNA sequence, determined that alternatively spliced RNAs and retroposed pseudogenes are incredibly abundant, and glimpsed the apparently huge number of non-coding RNAs that play significant roles in gene regulation. Ultimately, genome science is likely to provide comprehensive catalogs of these elements. However, the methods we have been using for most of the last 10 years will not yield even one complete open reading frame (ORF) for every gene--the first plateau on the long climb toward a comprehensive catalog. These strategies--sequencing randomly selected cDNA clones, aligning protein sequences identified in other organisms, sequencing more genomes, and manual curation--will have to be supplemented by large-scale amplification and sequencing of specific predicted mRNAs. The steady improvements in gene prediction that have occurred over the last 10 years have increased the efficacy of this approach and decreased its cost. In this Perspective, I review the state of gene prediction roughly 10 years ago, summarize the progress that has been made since, argue that the primary ORF identification methods we have relied on so far are inadequate, and recommend a path toward completing the Catalog of Protein Coding Genes, Version 1.0.

Animals↗

Genome annotation of a 1.5 Mb region of human chromosome 6q23 encompassing a quantitative trait locus for fetal hemoglobin expression in adults.

BACKGROUND: Heterocellular hereditary persistence of fetal hemoglobin (HPFH) is a common multifactorial trait characterized by a modest increase of fetal hemoglobin levels in adults. We previously localized a Quantitative Trait Locus for HPFH in an extensive Asian-Indian kindred to chromosome 6q23. As part of the strategy of positional cloning and a means towards identification of the specific genetic alteration in this family, a thorough annotation of the candidate interval based on a strategy of in silico / wet biology approach with comparative genomics was conducted. RESULTS: The ~1.5 Mb candidate region was shown to contain five protein-coding genes. We discovered a very large uncharacterized gene containing WD40 and SH3 domains (AHI1), and extended the annotation of four previously characterized genes (MYB, ALDH8A1, HBS1L and PDE7B). We also identified several genes that do not appear to be protein coding, and generated 17 kb of novel transcript sequence data from re-sequencing 97 EST clones. CONCLUSION: Detailed and thorough annotation of this 1.5 Mb interval in 6q confirms a high level of aberrant transcripts in testicular tissue. The candidate interval was shown to exhibit an extraordinary level of alternate splicing - 19 transcripts were identified for the 5 protein coding genes, but it appears that a significant portion (14/19) of these alternate transcripts did not have an open reading frame, hence their functional role is questionable. These transcripts may result from aberrant rather than regulated splicing.

3',5'-Cyclic-AMP Phosphodiesterases↗

Genomic annotation of 15,809 ESTs identified from pooled early gestation human eyes.

To complement cDNA libraries from the human eye at early gestation and to discover candidate genes associated with early ocular development, we used freshly dissected human eyeballs from week 9-14 of gestation to construct the early human fetal eye cDNA library. A total of 15,809 clones were isolated and sequenced from the unamplified and unnormalized library. We screened 11,246 good-quality ESTs, leading to the identification of 5,534 nonredundant clusters. Among them, 4,010 (72%) genes matched in the human protein database (Ensembl). The remaining 28% (1,524) corresponded to potentially novel or previously unidentified ESTs. We used BLASTX to compare our EST data with eight organisms and found common expression of a high portion of genes: Caenorhabditis briggsae (26%), Caenorhabditis elegans (27%), Anopheles gambiae (37%), Drosophila melanogaster (32%), Danio rerio (42%), Fugu rubripes (49%), Rattus norvegicusvalitus (52%), and Mus musculus (59%). Nevertheless, 48% (2,680 of 5,534) of the genes expressed in the early developing eye were not shared with current NEIBank human eye cDNA data. In addition, eight known retinal disease genes existed in our ESTs. Among them, six (COL11A1, BBS5, PDE6B, OAT, VMD2, and PGK1) were conserved among the genomes of other organisms, indicating that our annotated EST set provides not only a valuable resource for gene discovery and functional genomic analysis but also for phylogenetic analysis. Our foremost early gestation human eye cDNA library could provide detailed comparisons across species to identify physiological functions of genes and to elucidate evolutionary mechanisms.

Animals↗

Applications of DNA tiling arrays to experimental genome annotation and regulatory pathway discovery.

Microarrays have become a popular and important technology for surveying global patterns in gene expression and regulation. A number of innovative experiments have extended microarray applications beyond the measurement of mRNA expression levels, in order to uncover aspects of large-scale chromosome function and dynamics. This has been made possible due to the recent development of tiling arrays, where all non-repetitive DNA comprising a chromosome or locus is represented at various sequence resolutions. Since tiling arrays are designed to contain the entire DNA sequence without prior consultation of existing gene annotation, they enable the discovery of novel transcribed sequences and regulatory elements through the unbiased interrogation of genomic loci. The implementation of such methods for the global analysis of large eukaryotic genomes presents significant technical challenges. Nonetheless, tiling arrays are expected to become instrumental for the genome-wide identification and characterization of functional elements. Combined with computational methods to relate these data and map the complex interactions of transcriptional regulators, tiling array experiments can provide insight toward a more comprehensive understanding of fundamental molecular and cellular processes.

Gene Expression Profiling↗

Transcript mapping and genome annotation of ascidian mtDNA using EST data.

Mitochondrial transcripts of two ascidian species were reconstructed through sequence assembly of publicly available ESTs resembling mitochondrial DNA sequences (mt-ESTs). This strategy allowed us to analyze processing and mapping of the mitochondrial transcripts and to investigate the gene organization of a previously uncharacterized mitochondrial genome (mtDNA). This new strategy would greatly facilitate the sequencing and annotation of mtDNAs. In Ciona intestinalis, the assembled mt-ESTs covered 22 mitochondrial genes ( approximately 12,000 bp) and provided the partial sequence of the mtDNA and the prediction of its gene organization. Such sequences were confirmed by amplification and sequencing of the entire Ciona mtDNA. For Halocynthia roretzi, for which the mtDNA sequence was already available, the inferred mt transcripts allowed better definition of gene boundaries (16S rRNA, ND1, ATP6, and tRNA-Ser genes) and the identification of a new gene (an additional Phe-tRNA). In both species, polycistronic and immature transcripts, creation of stop codons by polyadenylation, tRNA signal processing, and rRNA transcript termination signals were identified, thus suggesting that the main features of mitochondrial transcripts are conserved in Chordata.

Animals↗

An agent-based system for re-annotation of genomes.

Genome annotation projects can produce incorrect results if they are based on obsolete data or inappropriate models. We have developed an automatic re-annotation system that uses agents to perform repetitive tasks and reports the results to the user. These tasks involve BLAST searches on biological databases (GenBank) and the use of detection tools (Genemark and Glimmer) to identify new open reading frames. Several agents execute these tools and combine their results to produce a list of open reading frames that is sent back to the user. Our goal was to reduce the manual work, executing most tasks automatically by computational tools. A prototype was implemented and validated using Mycoplasma pneumoniae and Haemophilus influenzae original annotated genomes. The results reported by the system identify most of new features present in the re-annotated versions of these genomes.

Computational Biology↗

Enhanced genome annotation using structural profiles in the program 3D-PSSM.

A method (three-dimensional position-specific scoring matrix, 3D-PSSM) to recognise remote protein sequence homologues is described. The method combines the power of multiple sequence profiles with knowledge of protein structure to provide enhanced recognition and thus functional assignment of newly sequenced genomes. The method uses structural alignments of homologous proteins of similar three-dimensional structure in the structural classification of proteins (SCOP) database to obtain a structural equivalence of residues. These equivalences are used to extend multiply aligned sequences obtained by standard sequence searches. The resulting large superfamily-based multiple alignment is converted into a PSSM. Combined with secondary structure matching and solvation potentials, 3D-PSSM can recognise structural and functional relationships beyond state-of-the-art sequence methods. In a cross-validated benchmark on 136 homologous relationships unambiguously undetectable by position-specific iterated basic local alignment search tool (PSI-Blast), 3D-PSSM can confidently assign 18 %. The method was applied to the remaining unassigned regions of the Mycoplasma genitalium genome and an additional 13 regions were assigned with 95 % confidence. 3D-PSSM is available to the community as a web server: http://www.bmm.icnet.uk/servers/3dpssm

Algorithms↗

The GAIA software framework for genome annotation.

We describe a software framework, GAIA, that supports semi-automated annotation of uncharacterized sequence data. The annotation framework incorporates annotation by data source integration, data analysis, and manual data entry. Components of the system include a configurable, open data analysis pipeline, a relational information storage manager, and Java-based graphical user interfaces. We discuss design decisions and tradeoffs in building such a system, and policies and strategies for producing consistent, uniform, high quality annotation.

Base Sequence↗