Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 955 records · Page 53Linked to original sources

Anopheles gambiae genome reannotation through synthesis of ab initio and comparative gene prediction algorithms.

BACKGROUND: Complete genome annotation is a necessary tool as Anopheles gambiae researchers probe the biology of this potent malaria vector. RESULTS: We reannotate the A. gambiae genome by synthesizing comparative and ab initio sets of predicted coding sequences (CDSs) into a single set using an exon-gene-union algorithm followed by an open-reading-frame-selection algorithm. The reannotation predicts 20,970 CDSs supported by at least two lines of evidence, and it lowers the proportion of CDSs lacking start and/or stop codons to only approximately 4%. The reannotated CDS set includes a set of 4,681 novel CDSs not represented in the Ensembl annotation but with EST support, and another set of 4,031 Ensembl-supported genes that undergo major structural and, therefore, probably functional changes in the reannotated set. The quality and accuracy of the reannotation was assessed by comparison with end sequences from 20,249 full-length cDNA clones, and evaluation of mass spectrometry peptide hit rates from an A. gambiae shotgun proteomic dataset confirms that the reannotated CDSs offer a high quality protein database for proteomics. We provide a functional proteomics annotation, ReAnoXcel, obtained by analysis of the new CDSs through the AnoXcel pipeline, which allows functional comparisons of the CDS sets within the same bioinformatic platform. CDS data are available for download. CONCLUSION: Comprehensive A. gambiae genome reannotation is achieved through a combination of comparative and ab initio gene prediction algorithms.

Algorithms↗

Functional screening for proapoptotic genes by reverse transfection cell array technology.

Application of mathematical algorithms to sequenced whole genomes revealed a large number of predicted genes, requiring functional assays for their characterization in a high-throughput manner. Here, we report on the development of a screening assay, which is based on reverse transfection of cellular arrays and subsequent analysis of cell morphology to identify novel proapoptotic genes. Expression plasmids containing full-length cDNAs were cotransfected with the reporter plasmid pEYFP to screen for apoptotic body formation, based on EYFP fluorescence. The assay was validated and applied to 382 human sequence-verified full-length open reading frames, most of them of unknown function. In this initial screening, proapoptotic effects could be demonstrated for 10 of these genes. For 6 of them apoptosis induction could be confirmed both by TUNEL assay and by FACS analysis of cells stained according to Nicoletti: 1 gene was not yet annotated for an apoptotic function (ST6GAL2), while 5 genes were without annotated function (FLJ20551, CXorf12, FAM105A, TMEM66, C19orf4). Our study demonstrates the potential of this method to characterize functionally genes of unknown function in a highly parallel format.

Apoptosis↗

A functional hierarchical organization of the protein sequence space.

BACKGROUND: It is a major challenge of computational biology to provide a comprehensive functional classification of all known proteins. Most existing methods seek recurrent patterns in known proteins based on manually-validated alignments of known protein families. Such methods can achieve high sensitivity, but are limited by the necessary manual labor. This makes our current view of the protein world incomplete and biased. This paper concerns ProtoNet, a automatic unsupervised global clustering system that generates a hierarchical tree of over 1,000,000 proteins, based solely on sequence similarity. RESULTS: In this paper we show that ProtoNet correctly captures functional and structural aspects of the protein world. Furthermore, a novel feature is an automatic procedure that reduces the tree to 12% its original size. This procedure utilizes only parameters intrinsic to the clustering process. Despite the substantial reduction in size, the system's predictive power concerning biological functions is hardly affected. We then carry out an automatic comparison with existing functional protein annotations. Consequently, 78% of the clusters in the compressed tree (5,300 clusters) get assigned a biological function with a high confidence. The clustering and compression processes are unsupervised, and robust. CONCLUSIONS: We present an automatically generated unbiased method that provides a hierarchical classification of all currently known proteins.

Algorithms↗

Get ready to GO! A biologist's guide to the Gene Ontology.

The Gene Ontology (GO) project provides a controlled vocabulary to facilitate high-quality functional gene annotation for all species. Genes in biological databases are linked to GO terms, allowing biologists to ask questions about gene function in a manner independent of species. This tutorial provides an introduction for biologists to the GO resources and covers three of the most common methods of querying GO: by individual gene, by gene function and by using a list of genes. [For the sake of brevity, the term 'gene' is used throughout this paper to refer to genes and their products (proteins and RNAs). GO annotations are always based on the characteristics of gene products, even though it may be the gene that is cited in the annotation.].

Abstracting and Indexing↗

Deep learning-based annotation of plant abiotic stress resistance genes for crops.

The declining costs of DNA sequencing have expanded genomic data, crucial for understanding plant abiotic stress responses and crop improvement. However, accurate gene annotation remains challenging. To address this limitation, we propose the PASRGA, a deep learning approach that leverages transfer learning and contrastive learning to annotate genes related to drought, salt, cold, and UV resistance. PASRGA achieves high F1-scores, area under the receiver operating characteristic (AUROC), area under the precision-recall curve (AUPRC), and Matthews correlation coefficient (MCC) in annotating stress resistance genes, significantly outperforming the general protein annotation model CLEAN, the plant phosphatase gene annotation model PF-NET, the top-ranked model in the CAFA5 challenge NetGO 4.0, and four traditional machine learning methods. Its effectiveness was further validated with a salt stress treatment experiment in Eutrema salsugineum. To facilitate crop breeding practices, we utilized PASRGA to annotate the genomes of 17 major crops. To improve accessibility and utility, we incorporated both manually curated and PASRGA-predicted gene data, together with the PASRGA tool, into the PlantASRG database (https://bioinfor.nefu.edu.cn/PlantASRG/). This comprehensive resource aims to support crop breeding initiatives and ensure food security.

Crops, Agricultural↗

iMOTdb--a comprehensive collection of spatially interacting motifs in proteins.

Realization of conserved residues that represent a protein family is crucial for clearer understanding of biological function as well as for the better recognition of additional members in sequence databases. Functionally important residues are recognized well due to their high degree of conservation in closely related sequences and are annotated in functional motif databases. Structural motifs are central to the integrity of the fold and require careful analysis for their identification. We report the availability of a database of spatially interacting motifs in single protein structures as well as those among distantly related protein structures that belong to a superfamily. Spatial interactions amongst conserved motifs are automatically measured using sequence similarity scores and distance calculations. Interactions between pairs of conserved motifs are described in the form of pseudoenergies. iMOTdb database provides information for 854,488 motifs corresponding to 60,849 protein structural domains and 22,648 protein structural entries.

Amino Acid Motifs↗

Improved techniques for the identification of pseudogenes.

MOTIVATION: Pseudogenes are the remnants of genomic sequences of genes which are no longer functional. They are frequent in most eukaryotic genomes, and an important resource for comparative genomics. However, pseudogenes are often mis-annotated as functional genes in sequence databases. Current methods for identifying pseudogenes include methods which rely on the presence of stop codons and frameshifts, as well as methods based on the ratio of non-silent to silent nucleotide substitution rates (dN/dS). A recent survey concluded that 50% of human pseudogenes have no detectable truncation in their pseudo-coding regions, indicating that the former methods lack sensitivity. The latter methods have been used to find sets of genes enriched for pseudogenes, but are not specific enough to accurately separate pseudogenes from expressed genes. RESULTS: We introduce a program called pseudogene inference from loss of constraint (PSILC) which incorporates novel methods for separating pseudogenes from functional genes. The methods calculate the log-odds score that evolution along the final branch of the gene tree to the query gene has been according to the following constraints: A neutral nucleotide model compared to a Pfam domain encoding model (PSILC(nuc/dom)); A protein coding model compared to a Pfam domain encoding model (PSILC(prot/dom)). Using the manual annotation of human chromosome 6, we show that both these methods result in a more accurate classification of pseudogenes than dN/dS when a Pfam domain alignment is available. AVAILABILITY: PSILC is available from http://www.sanger.ac.uk/Software/PSILC

Algorithms↗

Gramene, a tool for grass genomics.

Gramene (http://www.gramene.org) is a comparative genome mapping database for grasses and a community resource for rice (Oryza sativa). It combines a semi-automatically generated database of cereal genomic and expressed sequence tag sequences, genetic maps, map relations, and publications, with a curated database of rice mutants (genes and alleles), molecular markers, and proteins. Gramene curators read and extract detailed information from published sources, summarize that information in a structured format, and establish links to related objects both inside and outside the database, providing seamless connections between independent sources of information. Genetic, physical, and sequence-based maps of rice serve as the fundamental organizing units and provide a common denominator for moving across species and genera within the grass family. Comparative maps of rice, maize (Zea mays), sorghum (Sorghum bicolor), barley (Hordeum vulgare), wheat (Triticum aestivum), and oat (Avena sativa) are anchored by a set of curated correspondences. In addition to sequence-based mappings found in comparative maps and rice genome displays, Gramene makes extensive use of controlled vocabularies to describe specific biological attributes in ways that permit users to query those domains and make comparisons across taxonomic groups. Proteins are annotated for functional significance using gene ontology terms that have been adopted by numerous model species databases. Genetic variants including phenotypes are annotated using plant ontology terms common to all plants and trait ontology terms that are specific to rice. In this paper, we present a brief overview of the search tools available to the plant research community in Gramene.

Avena↗

A compendium of Caenorhabditis elegans regulatory transcription factors: a resource for mapping transcription regulatory networks.

BACKGROUND: Transcription regulatory networks are composed of interactions between transcription factors and their target genes. Whereas unicellular networks have been studied extensively, metazoan transcription regulatory networks remain largely unexplored. Caenorhabditis elegans provides a powerful model to study such metazoan networks because its genome is completely sequenced and many functional genomic tools are available. While C. elegans gene predictions have undergone continuous refinement, this is not true for the annotation of functional transcription factors. The comprehensive identification of transcription factors is essential for the systematic mapping of transcription regulatory networks because it enables the creation of physical transcription factor resources that can be used in assays to map interactions between transcription factors and their target genes. RESULTS: By computational searches and extensive manual curation, we have identified a compendium of 934 transcription factor genes (referred to as wTF2.0). We find that manual curation drastically reduces the number of both false positive and false negative transcription factor predictions. We discuss how transcription factor splice variants and dimer formation may affect the total number of functional transcription factors. In contrast to mouse transcription factor genes, we find that C. elegans transcription factor genes do not undergo significantly more splicing than other genes. This difference may contribute to differences in organism complexity. We identify candidate redundant worm transcription factor genes and orthologous worm and human transcription factor pairs. Finally, we discuss how wTF2.0 can be used together with physical transcription factor clone resources to facilitate the systematic mapping of C. elegans transcription regulatory networks. CONCLUSION: wTF2.0 provides a starting point to decipher the transcription regulatory networks that control metazoan development and function.

Animals↗

Characterization of 954 bovine full-CDS cDNA sequences.

BACKGROUND: Genome assemblies rely on the existence of transcript sequence to stitch together contigs, verify assembly of whole genome shotgun reads, and annotate genes. Functional genomics studies also rely on transcript sequence to create expression microarrays or interpret digital tag data produced by methods such as Serial Analysis of Gene Expression (SAGE). Transcript sequence can be predicted based on reconstruction from overlapping expressed sequence tags (EST) that are obtained by single-pass sequencing of random cDNA clones, but these reconstructions are prone to errors caused by alternative splice forms, transcripts from gene families with related sequences, and expressed pseudogenes. These errors confound genome assembly and annotation. The most useful transcript sequences are derived by complete insert sequencing of clones containing the entire length, or at least the full protein coding sequence (CDS) portion, of the source mRNA. While the bovine genome sequencing initiative is nearing completion, there is currently a paucity of bovine full-CDS mRNA and protein sequence data to support bovine genome assembly and functional genomics studies. Consequently, the production of high-quality bovine full-CDS cDNA sequences will enhance the bovine genome assembly and functional studies of bovine genes and gene products. The goal of this investigation was to identify and characterize the full-CDS sequences of bovine transcripts from clones identified in non-full-length enriched cDNA libraries. In contrast to several recent full-length cDNA investigations, these full-CDS cDNAs were selected, sequenced, and annotated without the benefit of the target organism's genomic sequence, by using comparison of bovine EST sequence to existing human mRNA to identify likely full-CDS clones for full-length insert cDNA (FLIC) sequencing. RESULTS: The predicted bovine protein lengths, 5' UTR lengths, and Kozak consensus sequences from 954 bovine FLIC sequences (bFLICs; average length 1713 nt, representing 762 distinct loci) are all consistent with previously sequenced mammalian full-length transcripts. CONCLUSION: In most cases, the bFLICs span the entire CDS of the genes, providing the basis for creating predicted bovine protein sequences to support proteomics and comparative evolutionary research as well as functional genomics and genome annotation. The results demonstrate the utility of the comparative approach in obtaining predicted protein sequences in other species.

5' Untranslated Regions↗

Arabidopsis thaliana full genome longmer microarrays: a powerful gene discovery tool for agriculture and forestry.

Sequenced plant genomes provide a large reservoir of known genes with potential for use in crop and tree improvement, but assignment of specific functions to annotated genes in sequenced plant genomes remains a challenge. Furthermore, most plant genes belong to families encoding proteins with related but distinct functions. In this commentary, we discuss our development of Arabidopsis spotted whole genome longmer oligonucleotide microarrays, and their use in global transcription profiling. We show that longmer array based transcriptome analysis in Arabidopsis can be used as an efficient and effective gene discovery and functional genomics tool, particularly for functional analyses of members of large gene families. We discuss experiments that focus on gene families involved in phenylpropanoid natural product biosynthesis and fiber differentiation. These analyses have helped to elucidate functions of individual gene family members, and have identified new candidate genes involved in fiber development and differentiation. Results obtained by these studies in Arabidopsis can be used as the basis for gene discovery in commercially important plants, and we have focused our attention on Populus trichocarpa (poplar), a species important in forestry and agroforestry for which complete genome sequence information is available.

Agriculture↗

Tackling non-canonical splicing in arrhythmogenic cardiomyopathy to reduce the uncertain significance variants burden.

BACKGROUND: Splice-altering variants (SAVs), particularly those outside canonical splice sites, are an underappreciated contributor to inherited cardiovascular diseases. In arrhythmogenic cardiomyopathy (ACM), these variants frequently remain classified as of uncertain significance (VUS) due to limited predictive power and lack of transcript-level evidence, constraining genetic yield and clinical management. Our study aimed to determine the functional impact of SAVs in ACM genes and refine their classification using ACMG/AMP and ClinGen SVI criteria. METHODS: SAVs identified in 200 ACM probands underwent SpliceAI prediction, GTEx cardiac exon-usage annotation, and functional assessment using pSPL3-based minigene assays. Aberrant transcripts were quantified using Percent Splicing Alteration (PSA). Segregation data and ACMG/AMP criteria refined by ClinGen SVI were applied to integrate functional and clinical evidence for classification. RESULTS: Aberrant splicing was confirmed in 9/20 variants (45%), including synonymous, missense, and non-canonical intronic changes. SpliceAI scores correlated strongly with PSA values (R²=0.86). Case-control burden testing revealed significant enrichment of splice-altering variants in DSP, DSG2, DSC2 and FLNC. Integrating predictive algorithms with experimental validation and segregation analysis markedly enhances reclassification of 16/20 variants (80%). CONCLUSION: Splicing defects beyond canonical sites significantly shape ACM genetic landscape. Integrating predictive models with experimental validation clarifies uncertain variants bridging the gap between genomic uncertainty and clinical decision-making.

Humans↗

Whole-plant growth stage ontology for angiosperms and its application in plant biology.

Plant growth stages are identified as distinct morphological landmarks in a continuous developmental process. The terms describing these developmental stages record the morphological appearance of the plant at a specific point in its life cycle. The widely differing morphology of plant species consequently gave rise to heterogeneous vocabularies describing growth and development. Each species or family specific community developed distinct terminologies for describing whole-plant growth stages. This semantic heterogeneity made it impossible to use growth stage description contained within plant biology databases to make meaningful computational comparisons. The Plant Ontology Consortium (http://www.plantontology.org) was founded to develop standard ontologies describing plant anatomical as well as growth and developmental stages that can be used for annotation of gene expression patterns and phenotypes of all flowering plants. In this article, we describe the development of a generic whole-plant growth stage ontology that describes the spatiotemporal stages of plant growth as a set of landmark events that progress from germination to senescence. This ontology represents a synthesis and integration of terms and concepts from a variety of species-specific vocabularies previously used for describing phenotypes and genomic information. It provides a common platform for annotating gene function and gene expression in relation to the developmental trajectory of a plant described at the organismal level. As proof of concept the Plant Ontology Consortium used the plant ontology growth stage ontology to annotate genes and phenotypes in plants with initial emphasis on those represented in The Arabidopsis Information Resource, Gramene database, and MaizeGDB.

Arabidopsis↗

Pseudo-messenger RNA: phantoms of the transcriptome.

The mammalian transcriptome harbours shadowy entities that resist classification and analysis. In analogy with pseudogenes, we define pseudo-messenger RNA to be RNA molecules that resemble protein-coding mRNA, but cannot encode full-length proteins owing to disruptions of the reading frame. Using a rigorous computational pipeline, which rules out sequencing errors, we identify 10,679 pseudo-messenger RNAs (approximately half of which are transposon-associated) among the 102,801 FANTOM3 mouse cDNAs: just over 10% of the FANTOM3 transcriptome. These comprise not only transcribed pseudogenes, but also disrupted splice variants of otherwise protein-coding genes. Some may encode truncated proteins, only a minority of which appear subject to nonsense-mediated decay. The presence of an excess of transcripts whose only disruptions are opal stop codons suggests that there are more selenoproteins than currently estimated. We also describe compensatory frameshifts, where a segment of the gene has changed frame but remains translatable. In summary, we survey a large class of non-standard but potentially functional transcripts that are likely to encode genetic information and effect biological processes in novel ways. Many of these transcripts do not correspond cleanly to any identifiable object in the genome, implying fundamental limits to the goal of annotating all functional elements at the genome sequence level.

Animals↗

Predicting functions from protein sequences--where are the bottlenecks?

The exponential growth of sequence data does not necessarily lead to an increase in knowledge about the functions of genes and their products. Prediction of function using comparative sequence analysis is extremely powerful but, if not performed appropriately, may also lead to the creation and propagation of assignment errors. While current homology detection methods can cope with the data flow, the identification, verification and annotation of functional features need to be drastically improved.

Amino Acid Sequence↗

A comprehensive BAC resource.

The Human Genome Project has generated extensive map and sequence data for a large number of Bacterial Artificial Chromosome (BAC) clones. In order to maximize the efficient use of the data and to minimize the redundant work for the research community, The Institute for Genomic Research (TIGR) comprehensive BAC resource (cBACr) (http://www.tigr.org/tdb/BacResource/BAC_resourc e_intro. html) was built as an expansion of the TIGR human BAC ends database. This resource collects, integrates and reports the information on library, maps, sequence, annotation and functions for each human and mouse BAC. The current database contains 635 016 human BACs and 265 617 mouse BACs that were characterized by various approaches, among which 22 705 human clones and 1000 mouse clones have sequence and annotation data.

Animals↗

Genome-wide high-throughput screens in functional genomics.

The availability of complete genome sequences from many organisms has yielded the ability to perform high-throughput, genome-wide screens of gene function. Within the past year, rapid advances have been made towards this goal in many major model systems, including yeast, worms, flies, and mammals. Yeast genome-wide screens have taken advantage of libraries of deletion strains, but RNA-interference has been used in other organisms to knockdown gene function. Examples of recent large-scale functional genetic screens include drug-target identification in yeast, regulators of fat accumulation in worms, growth and viability in flies, and proteasome-mediated degradation in mammalian cells. Within the next five years, such screens are likely to lead to annotation of function of most genes across multiple organisms. Integration of such data with other genomic approaches will extend our understanding of cellular networks.

Animals↗

High-resolution BAC-based map of the central portion of mouse chromosome 5.

The current strategy for sequencing the mouse genome involves the combination of a whole-genome shotgun approach with clone-based sequencing. High-resolution physical maps will provide a foundation for assembling contiguous segments of sequence. We have established a bacterial artificial chromosome (BAC)-based map of a 5-Mb region on mouse Chromosome 5, encompassing three gene families: receptor tyrosine kinases (PdgfraKit-Kdr), nonreceptor protein-tyrosine type kinases (Tec-Txk), and type-A receptors for the neurotransmitter GABA (Gabra2, Gabrb1, Gabrg1, and Gabra4). The construction of a BAC contig was initiated by hybridization screening the C57BL/6J (RPCI-23) BAC library, using known genes and sequence tagged sites (STSs). Additional overlapping clones were identified by searching the database of available restriction fingerprints for the RPCI-23 and RPCI-24 libraries. This effort resulted in the selection of >600 BAC clones, 251 kb of BAC-end sequences, and the placement of 40 known and/or predicted genes within this 5-Mb region. We use this high-resolution map to illustrate the integration of the BAC fingerprint map with a radiation-hybrid map via assembled expressed sequence tags (ESTs). From annotation of three representative BAC clones we demonstrate that up to 98% of the draft sequence for each contig could be ordered and oriented using known genes, BAC ends, consensus sequences for transcript assemblies, and comparisons with orthologous human sequence. For functional studies, annotation of sequence fragments as they are assembled into 50-200-kb stretches will be remarkably valuable.

Animals↗