Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,657 records · Page 92Linked to original sources

A collection of 11 800 single-copy Ds transposon insertion lines in Arabidopsis.

More than 10 000 transposon-tagged lines were constructed by using the Activator (Ac)/Dissociation (Ds) system in order to collect insertional mutants as a useful resource for functional genomics of Arabidopsis. The flanking sequences of the Ds element in the 11 800 independent lines were determined by high-throughput analysis using a semi-automated method. The sequence data allowed us to map the unique insertion site on the Arabidopsis genome in each line. The Ds element of 7566 lines is inserted in or close to coding regions, potentially affecting the function of 5031 of 25 000 Arabidopsis genes. Half of the lines have Ds insertions on chromosome 1 (Chr. 1), in which donor lines have a donor site. In the other half, the Ds insertions are distributed throughout the other four chromosomes. The intrachromosomal distribution of Ds insertions varies with the donor lines. We found that there are hot spots for Ds transposition near the ends of every chromosome, and we found some statistical preference for Ds insertion targets at the nucleotide level. On the basis of systematic analysis of the Ds insertion sites in the 11 800 lines, we propose the use of Ds-tagged lines with a single insertion in annotated genes for systematic analysis of phenotypes (phenome analysis) in functional genomics. We have opened a searchable database of the insertion-site sequences and mutated genes (http://rarge.gsc.riken.go.jp/) and are depositing these lines in the RIKEN BioResource Center as available resources (http://www.brc.riken.go.jp/Eng/).

Arabidopsis↗

The PEDANT genome database in 2005.

The PEDANT genome database (http://pedant.gsf.de) contains pre-computed bioinformatics analyses of publicly available genomes. Its main mission is to provide robust automatic annotation of the vast majority of amino acid sequences, which have not been subjected to in-depth manual curation by human experts in high-quality protein sequence databases. By design PEDANT annotation is genome-oriented, making it possible to explore genomic context of gene products, and evaluate functional and structural content of genomes using a category-based query mechanism. At present, the PEDANT database contains exhaustive annotation of over 1,240,000 proteins from 270 eubacterial, 23 archeal and 41 eukaryotic genomes.

Computational Biology↗

Bacterial genomes as new gene homes: the genealogy of ORFans in E. coli.

Differences in gene repertoire among bacterial genomes are usually ascribed to gene loss or to lateral gene transfer from unrelated cellular organisms. However, most bacteria contain large numbers of ORFans, that is, annotated genes that are restricted to a particular genome and that possess no known homologs. The uniqueness of ORFans within a genome has precluded the use of a comparative approach to examine their function and evolution. However, by identifying sequences unique to monophyletic groups at increasing phylogenetic depths, we can make direct comparisons of the characteristics of ORFans of different ages in the Escherichia coli genome, and establish their functional status and evolutionary rates. Relative to the genes ancestral to gamma-Proteobacteria and to those genes distributed sporadically in other prokaryotic species, ORFans in the E. coli lineage are short, A+T rich, and evolve quickly. Moreover, most encode functional proteins. Based on these features, ORFans are not attributable to errors in gene annotation, limitations of current databases, or to failure of methods for detecting homology. Rather, ORFans in the genomes of free-living microorganisms apparently derive from bacteriophage and occasionally become established by assuming roles in key cellular functions.

AT Rich Sequence↗

ZCURVE_V: a new self-training system for recognizing protein-coding genes in viral and phage genomes.

BACKGROUND: It necessary to use highly accurate and statistics-based systems for viral and phage genome annotations. The GeneMark systems for gene-finding in virus and phage genomes suffer from some basic drawbacks. This paper puts forward an alternative approach for viral and phage gene-finding to improve the quality of annotations, particularly for newly sequenced genomes. RESULTS: The new system ZCURVE_V has been run for 979 viral and 212 phage genomes, respectively, and satisfactory results are obtained. To have a fair comparison with the currently available software of similar function, GeneMark, a total of 30 viral genomes that have not been annotated by GeneMark are selected to be tested. Consequently, the average specificity of both systems is well matched, however the average sensitivity of ZCURVE_V for smaller viral genomes (< 100 kb), which constitute the main parts of viral genomes sequenced so far, is higher than that of GeneMark. Additionally, for the genome of Amsacta moorei entomopoxvirus, probably with the lowest genomic GC content among the sequenced organisms, the accuracy of ZCURVE_V is much better than that of GeneMark, because the later predicts hundreds of false-positive genes. ZCURVE_V is also used to analyze well-studied genomes, such as HIV-1, HBV and SARS-CoV. Accordingly, the performance of ZCURVE_V is generally better than that of GeneMark. Finally, ZCURVE_V may be downloaded and run locally, particularly facilitating its utilization, whereas GeneMark is not downloadable. Based on the above comparison, it is suggested that ZCURVE_V may serve as a preferred gene-finding tool for viral and phage genomes newly sequenced. However, it is also shown that the joint application of both systems, ZCURVE_V and GeneMark, leads to better gene-finding results. The system ZCURVE_V is freely available at: http://tubic.tju.edu.cn/Zcurve_V/. CONCLUSION: ZCURVE_V may serve as a preferred gene-finding tool used for viral and phage genomes, especially for anonymous viral and phage genomes newly sequenced.

Algorithms↗

LEGER: knowledge database and visualization tool for comparative genomics of pathogenic and non-pathogenic Listeria species.

Listeria species are ubiquitous in the environment and often contaminate foods because they grow under conditions used for food preservation. Listeria monocytogenes, the human and animal pathogen, causes Listeriosis, an infection with a high mortality rate in risk groups such as immune-compromised individuals. Furthermore, L.monocytogenes is a model organism for the study of intracellular bacterial pathogens. The publication of its genome sequence and that of the non-pathogenic species Listeria innocua initiated numerous comparative studies and efforts to sequence all species comprising the genus. The Proteome database LEGER (http://leger2.gbf.de/cgi-bin/expLeger.pl) was developed to support functional genome analyses by combining information obtained by applying bioinformatics methods and from public databases to improve the original annotations. LEGER offers three unique key features: (i) it is the first comprehensive information system focusing on the functional assignment of genes and proteins; (ii) integrated visualization tools, KEGG pathway and Genome Viewer, alleviate the functional exploration of complex data; and (iii) LEGER presents results of systematic post-genome studies, thus facilitating analyses combining computational and experimental results. Moreover, LEGER provides an unpublished membrane proteome analysis of L.innocua and in total visualizes experimentally validated information about the subcellular localizations of 789 different listerial proteins.

Bacterial Proteins↗

New friendly tools for users of ESTHER, the database of the alpha/beta-hydrolase fold superfamily of proteins.

The structural alpha/beta-hydrolase fold is characterized by a beta-sheet core of five to eight strands connected by alpha-helices to form a alpha/beta/alpha sandwich. The superfamily members, exemplified by the cholinesterases, diverged from a common ancestor into a number of hydrolytic enzymes displaying a wide range of substrate specificities, along with proteins with no recognized hydrolytic activity. In the enzymes, the catalytic triad residues are presented on loops of which one, the nucleophile elbow, is the most conserved feature of the fold. Of the other proteins, which all lack from one to all of the catalytic residues, some may simply be 'inactive' enzymes while others have been shown to be involved in heterologous surface recognition functions. The ESTHER (for esterases, alpha/beta-hydrolase enzymes and relatives) database (http://bioweb.ensam.inra.fr.esther) gathers and annotates all the published pieces of information (gene and protein sequences; biochemical, pharmacological, and structural data) related to the superfamily, and connects them together to provide the bases for studying structure-function relationships within the superfamily. The most recent developments of the database are presented.

Cholinesterases↗

GeneLibrarian: an effective gene-information summarization and visualization system.

BACKGROUND: Abundant information about gene products is stored in online searchable databases such as annotation or literature. To efficiently obtain and digest such information, there is a pressing need for automated information-summarization and functional-similarity clustering of genes. RESULTS: We have developed a novel method for semantic measurement of annotation and integrated it with a biomedical literature summarization system to establish a platform, GeneLibrarian, to provide users well-organized information about any specific group of genes (e.g. one cluster of genes from a microarray chip) they might be interested in. The GeneLibrarian generates a summarized viewgraph of candidate genes for a user based on his/her preference and delivers the desired background information effectively to the user. The summarization technique involves optimizing the text mining algorithm and Gene Ontology-based clustering method to enable the discovery of gene relations. CONCLUSION: GeneLibrarian is a Java-based web application that automates the process of retrieving critical information from the literature and expanding the number of potential genes for further analysis. This study concentrates on providing well organized information to users and we believe that will be useful in their researches. GeneLibrarian is available on http://gen.csie.ncku.edu.tw/GeneLibrarian/.

Algorithms↗

Experimental validation of data mined single nucleotide polymorphisms from several databases and consecutive dbSNP builds.

Rapid development in the annotation of human genetic variation has increased the numbers of single nucleotide polymorphisms (SNPs) in candidate genes by several orders of magnitude. The selection of both useful target SNPs for disease-gene association studies and SNPs associated with the treatment response is therefore an increasingly challenging task. We describe a workflow for selecting SNPs based on their putative function and frequency in candidate genes extracted from PubMed resources. The annotation of each SNP and its frequency in a Caucasian population was assessed in several databases. Approximately 4000 SNPs were identified from an initial 233 candidate genes. In a case study, we performed actual genotyping of 1030 of these SNPs in 213 genes and obtained 710 successfully genotyped SNPs. Using the flow-chart outlined here, only 87 SNPs were monomorphic (approximately 12%). This study reports the frequency of SNPs in a Caucasian population, selected in silico, using a candidate gene approach and validated by actually genotyping 193 individuals. The selected genotypes represent a valuable set of verified candidate SNPs for pharmacogenetic studies in Caucasian populations.

Breast Neoplasms↗

High-Resolution Chromosome-Level Genome Assembly and Annotation of Triplophysa stewarti, an Endemic Plateau Loach from the Qinghai-Tibet Plateau.

The bottom-dwelling fish Triplophysa stewarti, endemic to the Qinghai-Tibet Plateau, is a valuable model for studying high-altitude adaptation in aquatic ecosystems. However, the lack of a high-quality reference genome has hindered comparative genomic and evolutionary studies within this genus. Here, we present a chromosome-level genome assembly for T. stewarti, generated using PacBio HiFi long-read sequencing and Hi-C scaffolding. The 697.9&#x2009;Mb assembly is highly continuous (scaffold N50 of 253.58&#x2009;Mb) and encompasses 25 chromosomes, representing 92.65% of the genome. BUSCO analysis indicated a 98.4% completeness, supporting the high quality of the assembly. We annotated 28,009 protein-coding genes, with 97.04% being functionally assigned across multiple databases (NR, UniProt, KEGG, GO, Pfam and InterPro). Additionally, repetitive elements constituted 42.47% of the genome, and we identified 52,709 non-coding RNAs. This high-quality reference genome provides a fundamental resource for exploring the adaptive evolution, population structure, and conservation genetics of T. stewarti and related species on the Qinghai-Tibet Plateau.

Animals↗

PipTools: a computational toolkit to annotate and analyze pairwise comparisons of genomic sequences.

Sequence conservation between species is useful both for locating coding regions of genes and for identifying functional noncoding segments. Hence interspecies alignment of genomic sequences is an important computational technique. However, its utility is limited without extensive annotation. We describe a suite of software tools, PipTools, and related programs that facilitate the annotation of genes and putative regulatory elements in pairwise alignments. The alignment server PipMaker uses the output of these tools to display detailed information needed to interpret alignments. These programs are provided in a portable format for use on common desktop computers and both the toolkit and the PipMaker server can be found at our Web site (http://bio.cse.psu.edu/). We illustrate the utility of the toolkit using annotation of a pairwise comparison of the mouse MHC class II and class III regions with orthologous human sequences and subsequently identify conserved, noncoding sequences that are DNase I hypersensitive sites in chromatin of mouse cells.

Animals↗

SAWTED: structure assignment with text description--enhanced detection of remote homologues with automated SWISS-PROT annotation comparisons.

MOTIVATION: Sequence database search methods often identify putative sub-threshold hits of known function or structure for a given query sequence. It is widespread practice to filter these hits by hand using knowledge of function and other factors; to the expert, some hits may appear more sensible than others. SAWTED (Structure Assignment With Text Description) is an automated solution to this post-filtering problem which will be applicable to large scale genome assignments. RESULTS: A standard document comparison algorithm is applied to text descriptions extracted from SWISS-PROT annotations. The added value of SAWTED in combination with PSI-BLAST has been shown with a benchmark of difficult remote homologues taken from the SCOP structure database. AVAILABILITY: A WAWTED PSI-BLAST Web server is available to perform sensitive searches against the protein structure database (http://www.bmm.icnet.uk/servers/sawted). CONTACT: R.MacCallum@icrf.icnet.uk

Algorithms↗

Automatic annotation of eukaryotic genes, pseudogenes and promoters.

BACKGROUND: The ENCODE gene prediction workshop (EGASP) has been organized to evaluate how well state-of-the-art automatic gene finding methods are able to reproduce the manual and experimental gene annotation of the human genome. We have used Softberry gene finding software to predict genes, pseudogenes and promoters in 44 selected ENCODE sequences representing approximately 1% (30 Mb) of the human genome. Predictions of gene finding programs were evaluated in terms of their ability to reproduce the ENCODE-HAVANA annotation. RESULTS: The Fgenesh++ gene prediction pipeline can identify 91% of coding nucleotides with a specificity of 90%. Our automatic pseudogene finder (PSF program) found 90% of the manually annotated pseudogenes and some new ones. The Fprom promoter prediction program identifies 80% of TATA promoters sequences with one false positive prediction per 2,000 base-pairs (bp) and 50% of TATA-less promoters with one false positive prediction per 650 bp. It can be used to identify transcription start sites upstream of annotated coding parts of genes found by gene prediction software. CONCLUSION: We review our software and underlying methods for identifying these three important structural and functional genome components and discuss the accuracy of predictions, recent advances and open problems in annotating genomic sequences. We have demonstrated that our methods can be effectively used for initial automatic annotation of the eukaryotic genome.

Animals↗

Design of a system for combined analysis of microarray-based gene expression and FlyBase-derived annotation in Drosophila.

In recent years, researchers began to utilize both the experimental microarray data and the annotation data from FlyBase to identify Drosophila genes that might encode certain functions. So far, they have to manually combine data from both the microarray experiment and the FlyBase, which is a slow and tedious process. The goal of this research is to construct a flexible relational database system that integrates the microarray data with annotation information from FlyBase.

Animals↗

Comparative sequence analysis of Sordaria macrospora and Neurospora crassa as a means to improve genome annotation.

One of the most challenging parts of large scale sequencing projects is the identification of functional elements encoded in a genome. Recently, studies of genomes of up to six different Saccharomyces species have demonstrated that a comparative analysis of genome sequences from closely related species is a powerful approach to identify open reading frames and other functional regions within genomes [Science 301 (2003) 71, Nature 423 (2003) 241]. Here, we present a comparison of selected sequences from Sordaria macrospora to their corresponding Neurospora crassa orthologous regions. Our analysis indicates that due to the high degree of sequence similarity and conservation of overall genomic organization, S. macrospora sequence information can be used to simplify the annotation of the N. crassa genome.

Base Sequence↗

Hepatitis C databases, principles and utility to researchers.

Part of the effort to develop hepatitis C-specific drugs a nd vaccines is the study of genetic variability of allpublicly available HCV sequences. Three HCV databases are currently available to aid this effort and to provide additional insight into the basic biology, immunology, and evolution of the virus. The Japanese HCV database (http://s2as02.genes.nig.ac.jp) gives access to a genomic mapping of sequences as well as their phylogenetic relationships. The European HCV database (http://euhcvdb.ibcp.fr) offers access to a computer-annotated set of sequences and molecular models of HCV proteins and focuses on protein sequence, structure and function analysis. The HCV database at the Los Alamos National Laboratory in the United States (http://hcv.lanl.gov) provides access to a manually annotated sequence database and a database of immunological epitopes which contains concise descriptions of experimental results. In this paper, we briefly describe each of these databases and their associated websites and tools, and give some examples of their use in furthering HCV research.

Biomedical Research↗

ASAP: amplification, sequencing & annotation of plastomes.

BACKGROUND: Availability of DNA sequence information is vital for pursuing structural, functional and comparative genomics studies in plastids. Traditionally, the first step in mining the valuable information within a chloroplast genome requires sequencing a chloroplast plasmid library or BAC clones. These activities involve complicated preparatory procedures like chloroplast DNA isolation or identification of the appropriate BAC clones to be sequenced. Rolling circle amplification (RCA) is being used currently to amplify the chloroplast genome from purified chloroplast DNA and the resulting products are sheared and cloned prior to sequencing. Herein we present a universal high-throughput, rapid PCR-based technique to amplify, sequence and assemble plastid genome sequence from diverse species in a short time and at reasonable cost from total plant DNA, using the large inverted repeat region from strawberry and peach as proof of concept. The method exploits the highly conserved coding regions or intergenic regions of plastid genes. Using an informatics approach, chloroplast DNA sequence information from 5 available eudicot plastomes was aligned to identify the most conserved regions. Cognate primer pairs were then designed to generate approximately 1 - 1.2 kb overlapping amplicons from the inverted repeat region in 14 diverse genera. RESULTS: 100% coverage of the inverted repeat region was obtained from Arabidopsis, tobacco, orange, strawberry, peach, lettuce, tomato and Amaranthus. Over 80% coverage was obtained from distant species, including Ginkgo, loblolly pine and Equisetum. Sequence from the inverted repeat region of strawberry and peach plastome was obtained, annotated and analyzed. Additionally, a polymorphic region identified from gel electrophoresis was sequenced from tomato and Amaranthus. Sequence analysis revealed large deletions in these species relative to tobacco plastome thus exhibiting the utility of this method for structural and comparative genomics studies. CONCLUSION: This simple, inexpensive method now allows immediate access to plastid sequence, increasing experimental throughput and serving generally as a universal platform for plastid genome characterization. The method applies well to whole genome studies and speeds assessment of variability across species, making it a useful tool in plastid structural genomics.

Base Sequence↗

Receptor-defined targeting of a genomically unique melanoma-enriched noncanonical antigen.

Effective T cell-based immunotherapies require functional receptors that can be engineered and redeployed to recognize tumor-restricted antigens. Noncanonical peptides arising from transcription outside annotated protein-coding regions expand the antigenic landscape of cancer; however, systematic strategies to biologically prioritize and functionally validate such targets remain underdeveloped. Here, we integrated de novo transcript analysis, exon-resolved quantification, RNA in situ hybridization, and immunopeptidomics to identify melanoma-associated noncanonical transcripts and advance candidates through receptor-level validation. Among three recurrent melanoma-associated transcripts, EVA003 emerged as a lead target based on its distinct repeat-enriched genomic architecture, consistent tumor-enriched exon-level expression across independent datasets, and a genomically unique immunogenic core sequence. We demonstrate endogenous presentation of EVA003-derived peptides on HLA-A*03:01 and detect specific reactivity in patient-derived tumor-infiltrating lymphocytes. Single-cell transcriptomic profiling identified a dominant peptide-reactive clonotype, enabling isolation of a naturally occurring T cell receptor. Transfer of this receptor into healthy donor T cells conferred antigen-dependent activation and cytotoxicity against both peptide-pulsed targets and melanoma cells expressing EVA003 endogenously. Together, these findings establish a biologically informed strategy for prioritizing noncanonical tumor antigens and demonstrate that genomically unique, tumor-enriched noncanonical peptides can be presented to molecularly defined receptors capable of mediating cancer cell killing. These findings support the integration of prioritized noncanonical antigens into engineered T cell therapeutic strategies.

Humans↗

Genome-wide prediction and analysis of function-specific transcription factor binding sites.

DNA-binding transcription factors play a central role in transcription regulation, and the annotation of transcription-factor binding sites in upstream regions of human genes is essential for building a genome-wide regulatory network. We describe methodology to accurately predict the transcription-factor binding sites in the proximal-promoter region of function-specific genes. In order to increase the accuracy of transcription factor binding-site prediction, we rely on recent genome sequence data, known transcription factor binding-site matrices, and Gene Ontology biological-function-based gene classification. Using TRANSFAC position-frequency matrices, we detected individual and cooperating transcription-factor binding sites in proximal promoters of ENSEMBL annotated human genes. We used the over representation of detected binding sites in the proximal promoters as compared to the second exons to control specificity. We confirmed the majority of transcription-factor binding sites predicted in proximal promoters of immune-response genes with evidence from existing literature. We validated the predicted cooperation between transcription factors NF-kappa B and IRF in the regulation of gene expression with microarray transcript profiling data and literature-derived protein-protein interaction network. We also identified over-represented individual and pairs of transcription-factor binding sites in the proximal promoters of each Gene Ontology biological-process gene group. Our tools and analysis provide a new resource for deciphering transcription regulation in different biological paradigms.

Computational Biology↗