Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,459 records · Page 81Linked to original sources

ASAP: the Alternative Splicing Annotation Project.

Recently, genomics analyses have demonstrated that alternative splicing is widespread in mammalian genomes (30-60% of genes reported to have multiple isoforms), and may be one of their most important mechanisms of functional regulation. However, by comparison with other genomics data such as genome annotation, SNPs, or gene expression, there exists relatively little database infrastructure for the study of alternative splicing. We have constructed an online database ASAP (the Alternative Splicing Annotation Project) for biologists to access and mine the enormous wealth of alternative splicing information coming from genomics and proteomics. ASAP is based on genome-wide analyses of alternative splicing in human (30 793 alternative splice relationships found) from detailed alignment of expressed sequences onto the genomic sequence. ASAP provides precise gene exon-intron structure, alternative splicing, tissue specificity of alternative splice forms, and protein isoform sequences resulting from alternative splicing. Moreover, it can help biologists design probe sequences for distinguishing specific mRNA isoforms. ASAP is intended to be a community resource for collaborative annotation of alternative splice forms, their regulation, and biological functions. The URL for ASAP is http://www.bioinformatics.ucla.edu/ASAP.

Alternative Splicing↗

Gene expression analyses of Arabidopsis chromosome 2 using a genomic DNA amplicon microarray.

The gene predictions and accompanying functional assignments resulting from the sequencing and annotation of a genome represent hypotheses that can be tested and used to develop a more complete understanding of the organism and its biology. In the model plant Arabidopsis thaliana, we developed a novel approach to constructing whole-genome microarrays based on PCR amplification of the 3' ends of each predicted gene from genomic DNA, and constructed an array representing more than 94% of the predicted genes and pseudogenes on chromosome 2. With this array, we examined various tissues and physiological conditions, providing expression-based validation for 84% of the gene predictions and providing clues as to the functions of many predicted genes. Further, by examining the distribution of expression along the physical chromosome, we were able to identify a region of repressed transcription that may represent a previously undescribed heterochromatic region.

Arabidopsis↗

Specialized microbial databases for inductive exploration of microbial genome sequences.

BACKGROUND: The enormous amount of genome sequence data asks for user-oriented databases to manage sequences and annotations. Queries must include search tools permitting function identification through exploration of related objects. METHODS: The GenoList package for collecting and mining microbial genome databases has been rewritten using MySQL as the database management system. Functions that were not available in MySQL, such as nested subquery, have been implemented. RESULTS: Inductive reasoning in the study of genomes starts from "islands of knowledge", centered around genes with some known background. With this concept of "neighborhood" in mind, a modified version of the GenoList structure has been used for organizing sequence data from prokaryotic genomes of particular interest in China. GenoChore http://bioinfo.hku.hk/genochore.html, a set of 17 specialized end-user-oriented microbial databases (including one instance of Microsporidia, Encephalitozoon cuniculi, a member of Eukarya) has been made publicly available. These databases allow the user to browse genome sequence and annotation data using standard queries. In addition they provide a weekly update of searches against the world-wide protein sequences data libraries, allowing one to monitor annotation updates on genes of interest. Finally, they allow users to search for patterns in DNA or protein sequences, taking into account a clustering of genes into formal operons, as well as providing extra facilities to query sequences using predefined sequence patterns. CONCLUSION: This growing set of specialized microbial databases organize data created by the first Chinese bacterial genome programs (ThermaList, Thermoanaerobacter tencongensis, LeptoList, with two different genomes of Leptospira interrogans and SepiList, Staphylococcus epidermidis) associated to related organisms for comparison.

Algorithms↗

Chromosome-level Genome Assembly of the Halophytic Turfgrass Zoysia macrostachya.

Zoysia macrostachya Franch. & Sav. is a halophytic perennial turfgrass in the Poaceae family, commonly found in the coastal regions of Korea, Japan, and East Asia. Z. macrostachya thrives in high-salinity environments, making it an excellent model for studying abiotic stress resilience. In this study, we present a chromosome-level genome assembly of Z. macrostachya, constructed using Oxford Nanopore long reads, Illumina short reads, and Omni-C sequencing data. The assembly spans 329.78 Mb across 20 chromosomes, with a scaffold N50 of 19.24 Mb, and includes complete telomeric sequences at both ends. The assembly showed 97.8% complete BUSCOs, indicating high genome completeness. Repeat element and gene annotation identified 44.03% of the genome as repetitive elements and 33,474 protein-coding genes. The gene annotation showed 97.1% complete BUSCOs and 86.92% functionally characterized genes. Macrosynteny analysis highlighted highly collinear relationships with related species, providing a foundational understanding of the Z. macrostachya genomic structure. This high-quality genome serves as a valuable resource for advancing salinity tolerance research and improving the genetic diversity of Zoysia species.

Genome, Plant↗

Phylogenetic analysis of general bacterial porins: a phylogenomic case study.

Bacterial porin proteins allow for the selective movement of hydrophilic solutes through the outer membrane of Gram-negative bacteria. The purpose of this study was to clarify the evolutionary relationships among the Type 1 general bacterial porins (GBPs), a porin protein subfamily that includes outer membrane proteins ompC and ompF among others. Specifically, we investigated the potential utility of phylogenetic analysis for refining poorly annotated or mis-annotated protein sequences in databases, and for characterizing new functionally distinct groups of porin proteins. Preliminary phylogenetic analysis of sequences obtained from GenBank indicated that many of these sequences were incompletely or even incorrectly annotated. Using a well-curated set of porins classified via comparative genomics, we applied recently developed bayesian phylogenetic methods for protein sequence analysis to determine the relationships among the Type 1 GBPs. Our analysis found that the major GBP classes (ompC, phoE, nmpC and ompN) formed strongly supported monophyletic groups, with the exception of ompF, which split into two distinct clades. The relationships of the GBP groups to one another had less statistical support, except for the relationships of ompC and ompN sequences, which were strongly supported as sister groups. A phylogenetic analysis comparing the relationships of the GenBank GBP sequences to the correctly annotated set of GBPs identified a large number of previously unclassified and mis-annotated GBPs. Given these promising results, we developed a tree-parsing algorithm for automated phylogenetic annotation and tested it with GenBank sequences. Our algorithm was able to automatically classify 30 unidentified and 15 mis-annotated GBPs out of 78 sequences. Altogether, our results support the potential for phylogenomics to increase the accuracy of sequence annotations.

Algorithms↗

Comparative analysis of the expressed genome of the infective juvenile entomopathogenic nematode, Heterorhabditis bacteriophora.

We report the first cDNA-sequencing project of the entomopathogenic nematode, Heterorhabditis bacteriophora. A total of 1246 expressed sequence tags (ESTs) were generated by random sequencing of clones from a cDNA library of the infective juvenile stage. The ESTs were annotated resulting in 1072 useful ESTs that were categorized into functional categories according to Kyoto Encyclopedia of Genes and Genomes. Approximately 459 of 1072 ESTs (43%) had significant similarities to annotated sequences in GenBank. Of these, 417 had significant similarities to the free-living nematode Caenorhanditis elegans proteins. Most ESTs (18%) belonged to the genetic information processing category followed by metabolism (15% ESTs) and environmental information processing (15%) pathways. Several interesting ESTs were found that may have roles in the infectivity and survival of infective juveniles. These included proteases, dauer pathway genes (akt-1, pdk-1 & daf-7) and aging and stress resistance genes such as superoxide dismutase (sod-4), heat shock genes (hsp-4 & hsp-6), and eat genes, and signaling proteins like G-protein coupled receptors, regulators of G-protein signaling (rgs), and serine/threonine kinases. Other interesting ESTs include systemic RNAi defective protein (sid-1), ribonuclease III family members (rnh-2 &rnc) and transposase gene (Tc3A). About 67% of the ESTs did not find matches in any of the searched databases suggesting potentially novel genes in this enomopathogenic nematode. Note: Sequences described in this paper have been deposited in Genbank under the accessions DN 152655-DN 152999, and DN 153000-DN 153726.

3-Phosphoinositide-Dependent Protein Kinases↗

Pathways database system: an integrated system for biological pathways.

MOTIVATION: During the next phase of the Human Genome Project, research will focus on functional studies of attributing functions to genes, their regulatory elements, and other DNA sequences. To facilitate the use of genomic information in such studies, a new modeling perspective is needed to examine and study genome sequences in the context of many kinds of biological information. Pathways are the logical format for modeling and presenting such information in a manner that is familiar to biological researchers. RESULTS: In this paper we present an integrated system, called Pathways Database System, with a set of software tools for modeling, storing, analyzing, visualizing, and querying biological pathways data at different levels of genetic, molecular, biochemical and organismal detail. The novel features of the system include: (a) genomic information integrated with other biological data and presented from a pathway, rather than from the DNA sequence, perspective; (b) design for biologists who are possibly unfamiliar with genomics, but whose research is essential for annotating gene and genome sequences with biological functions; (c) database design, implementation and graphical tools which enable users to visualize pathways data in multiple abstraction levels, and to pose predetermined queries; and (d) an implementation that allows for web(XML)-based dissemination of query outputs (i.e. pathways data) to researchers in the community, giving them control on the use of pathways data. AVAILABILITY: Available on request from the authors.

Database Management Systems↗

The long hard road to a completed Candida albicans genome.

After almost a decade of work, the sequencing, assembly, and annotation of the genome of the fungal pathogen Candida albicans is finally close at hand. This review covers the early history of the C. albicans genome project, from the release of early assemblies that provided the impetus for an explosion in functional genomics research, to a community-based annotation and a preview of the work that was necessary for the production of a final genome assembly.

Base Sequence↗

A high-resolution map of transcription in the yeast genome.

There is abundant transcription from eukaryotic genomes unaccounted for by protein coding genes. A high-resolution genome-wide survey of transcription in a well annotated genome will help relate transcriptional complexity to function. By quantifying RNA expression on both strands of the complete genome of Saccharomyces cerevisiae using a high-density oligonucleotide tiling array, this study identifies the boundary, structure, and level of coding and noncoding transcripts. A total of 85% of the genome is expressed in rich media. Apart from expected transcripts, we found operon-like transcripts, transcripts from neighboring genes not separated by intergenic regions, and genes with complex transcriptional architecture where different parts of the same gene are expressed at different levels. We mapped the positions of 3' and 5' UTRs of coding genes and identified hundreds of RNA transcripts distinct from annotated genes. These nonannotated transcripts, on average, have lower sequence conservation and lower rates of deletion phenotype than protein coding genes. Many other transcripts overlap known genes in antisense orientation, and for these pairs global correlations were discovered: UTR lengths correlated with gene function, localization, and requirements for regulation; antisense transcripts overlapped 3' UTRs more than 5' UTRs; UTRs with overlapping antisense tended to be longer; and the presence of antisense associated with gene function. These findings may suggest a regulatory role of antisense transcription in S. cerevisiae. Moreover, the data show that even this well studied genome has transcriptional complexity far beyond current annotation.

5' Untranslated Regions↗

Functional clues for hypothetical proteins based on genomic context analysis in prokaryotes.

Three integrated genomic context methods were used to annotate uncharacterized proteins in 102 bacterial genomes. Of 7853 orthologous groups with unknown function containing 45,110 proteins, 1738 groups could be linked to functionally associated partners. In many cases, those partners are uncharacterized themselves (hinting at newly identified modules) or have been described in general terms only. However, we were able to assign pathways, cellular processes or physical complexes for 273 groups (encompassing 3624 previously functionally uncharacterized proteins).

Bacterial Proteins↗

miRBase: microRNA sequences, targets and gene nomenclature.

The miRBase database aims to provide integrated interfaces to comprehensive microRNA sequence data, annotation and predicted gene targets. miRBase takes over functionality from the microRNA Registry and fulfils three main roles: the miRBase Registry acts as an independent arbiter of microRNA gene nomenclature, assigning names prior to publication of novel miRNA sequences. miRBase Sequences is the primary online repository for miRNA sequence data and annotation. miRBase Targets is a comprehensive new database of predicted miRNA target genes. miRBase is available at http://microrna.sanger.ac.uk/.

Animals↗

Cytogenetic and molecular characterization of heterochromatin gene models in Drosophila melanogaster.

In the past decade, genome-sequencing projects have yielded a great amount of information on DNA sequences in several organisms. The release of the Drosophila melanogaster heterochromatin sequence by the Drosophila Heterochromatin Genome Project (DHGP) has greatly facilitated studies of mapping, molecular organization, and function of genes located in pericentromeric heterochromatin. Surprisingly, genome annotation has predicted at least 450 heterochromatic gene models, a figure 10-fold above that defined by genetic analysis. To gain further insight into the locations and functions of D. melanogaster heterochromatic genes and genome organization, we have FISH mapped 41 gene models relative to the stained bands of mitotic chromosomes and the proximal divisions of polytene chromosomes. These genes are contained in eight large scaffolds, which together account for approximately 1.4 Mb of heterochromatic DNA sequence. Moreover, developmental Northern analysis showed that the expression of 15 heterochromatic gene models tested is similar to that of the vital heterochromatic gene Nipped-A, in that it is not limited to specific stages, but is present throughout all development, despite its location in a supposedly "silent" region of the genome. This result is consistent with the idea that genes resident in heterochromatin can encode essential functions.

Animals↗

AMIGOS: a method for the inspection of genomic organisation or structure and its application to characterise conserved gene arrangements.

In order to identify and to characterise gene clusters conserved in microbial genomes, the algorithm AMIGOS was developed. It is based on a categorisation of genes using a predefined set of gene functions (GFs). After the categorisation of all genes of a genome and based on their location on a replicon, distances between GFs were determined and stored in genome-specific matrices. These matrices were used to identify GF clusters like those strictly conserved in 13 archaeal, in 47 bacterial genomes and in the combination of the sets. Within the combined set of these 60 microbial genomes, there exist only two strictly conserved clusters harbouring two ribosomal genes each, namely those for L4, L23 and L22, L29. In order to characterise less strictly conserved GF clusters, content of genomes i.e. matrices were analysed pairwise. Resulting clusters were merged to (meta-) clusters if their content overlapped. A scoring system named cons(CL) was developed. It quantifies conservedness of cluster membership for individual GFs. For the genome of Escherichia coli it was shown that a grouping of cluster elements on cons(CL) values dissected the clusters into smaller sets. These sets were frequently overlapped by known transcriptional units (TUs). This finding justifies the usage of cons(CL) scores to predict TU membership of genes. In addition, cons(CL) values provide a sound basis for non-homologous gene annotation. Based on cons(CL) values, examples of conserved clusters containing annotated genes and single ones with unknown function are given.

Algorithms↗

Gene annotation: prediction and testing.

Fifty years after the publication of DNA structure, the whole human genome sequence will be officially finished. This achievement marks the beginning of the task to catalogue every human gene and identify each of their function expression patterns. Currently, researchers estimate that there are about 30,000 human genes and approximately 70% of these can be automatically predicted using a combination of ab initio and similarity-based programs. However, to experimentally investigate every gene's function, the research community requires a high-quality annotation of alternative splicing, pseudogenes, and promoter regions that can only be provided by manual intervention. Manual curation of the human genome will be a long-term project as experimental data are continually produced to confirm or refine the predictions, and new features such as noncoding RNAs and enhancers have not been fully identified. Such a highly curated human gene-set made publicly available will be a great asset for the experimental community and for future comparative genome projects.

Alternative Splicing↗

Disruption of GAD1 protein architecture by a novel missense variant in a consanguineous family with autosomal recessive intellectual disability.

BACKGROUND: Intellectual disability represents a heterogeneous group of neurodevelopmental disorders marked by significant impairments in intellectual functioning and adaptive behavior. Among the various causes, genetic factors play a major role, with autosomal recessive intellectual disability (ARID) constituting a genetically diverse subgroup. ARID is prevalent in consanguineous families and arises from homozygous mutations that disrupt critical genes involved in brain development and function. OBJECTIVE: This study aimed to identify disease-causing genetic variants responsible for ARID in a consanguineous Pakistani family and to evaluate the structural and functional impact of a novel variant identified in GAD1 through protein modeling. METHODS: A consanguineous family affected with intellectual disability was enrolled. Whole-exome sequencing was performed on an affected individual, followed by bioinformatics analysis including alignment to the GRCh38 reference genome, variant calling, and annotation. Variants were filtered based on rarity, predicted functional impact, and autosomal recessive inheritance pattern. Candidate variants were validated and assessed by Sanger sequencing and segregation analysis. Protein modeling was performed to evaluate the structural impact of the identified variant. RESULTS: A novel homozygous missense variant NM_000817:c.1700G>A;p.Arg567Gln in GAD1 was identified. Segregation analysis confirmed co-segregation of the variant with the affected phenotype. Protein modeling suggested that the variant may disrupt GAD1 enzymatic function involved in gamma-aminobutyric acid synthesis. CONCLUSION: This study emphasizes the significance of genetic investigation in familial cases and the crucial role that GAD1 mutations play in neurodevelopmental disorders with intellectual disability. The results advance the knowledge of molecular causes of ARID and broaden the mutational range.

Pakistani↗

Recent developments in analytical and functional protein microarrays.

In recent years, the genomes of many different organisms have been fully sequenced and annotated. As a consequence of this information, a number of methods have emerged to study the function of many genes and proteins in parallel. One recent approach for the large-scale analysis of proteins is the use of protein microarrays in which hundreds to thousands of proteins are arrayed and assayed simultaneously. Protein arrays can be used for assessing protein levels and following disease markers, identifying biochemical activities, analyzing post-translational modifications, building interaction networks, and for drug discovery and development. In this review, we discuss the construction of different types of protein arrays, and their numerous and diverse applications.

Antibodies↗

Sequence, annotation, and analysis of synteny between rice chromosome 3 and diverged grass species.

Rice (Oryza sativa L.) chromosome 3 is evolutionarily conserved across the cultivated cereals and shares large blocks of synteny with maize and sorghum, which diverged from rice more than 50 million years ago. To begin to completely understand this chromosome, we sequenced, finished, and annotated 36.1 Mb ( approximately 97%) from O. sativa subsp. japonica cv Nipponbare. Annotation features of the chromosome include 5915 genes, of which 913 are related to transposable elements. A putative function could be assigned to 3064 genes, with another 757 genes annotated as expressed, leaving 2094 that encode hypothetical proteins. Similarity searches against the proteome of Arabidopsis thaliana revealed putative homologs for 67% of the chromosome 3 proteins. Further searches of a nonredundant amino acid database, the Pfam domain database, plant Expressed Sequence Tags, and genomic assemblies from sorghum and maize revealed only 853 nontransposable element related proteins from chromosome 3 that lacked similarity to other known sequences. Interestingly, 426 of these have a paralog within the rice genome. A comparative physical map of the wild progenitor species, Oryza nivara, with japonica chromosome 3 revealed a high degree of sequence identity and synteny between these two species, which diverged approximately 10,000 years ago. Although no major rearrangements were detected, the deduced size of the O. nivara chromosome 3 was 21% smaller than that of japonica. Synteny between rice and other cereals using an integrated maize physical map and wheat genetic map was strikingly high, further supporting the use of rice and, in particular, chromosome 3, as a model for comparative studies among the cereals.

Arabidopsis↗

From gene networks to brain networks.

The brain's structural organization is so complex that 2,500 years of analysis leaves pervasive uncertainty about (i) the identity of its basic parts (regions with their neuronal cell types and pathways interconnecting them), (ii) nomenclature, (iii) systematic classification of the parts with respect to topographic relationships and functional systems and (iv) the reliability of the connectional data itself. Here we present a prototype knowledge management system (http://brancusi.usc.edu/bkms/) for analyzing the architecture of brain networks in a systematic, interactive and extendable way. It supports alternative interpretations and models, is based on fully referenced and annotated data and can interact with genomic and functional knowledge management systems through web services protocols.

Animals↗