Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,621 records · Page 90Linked to original sources

A new member of plant CS-lyases. A cystine lyase from Arabidopsis thaliana.

Cystine lyases catalyze the breakdown of l-cystine to thiocysteine, pyruvate, and ammonia. Until now there are no reports of the identification of a plant cystine lyase at a molecular level, and it is not clear what biological role this class of enzymes have in plants. A cystine lyase was isolated from Brassica oleracea (L.), and partial amino acid sequencing allowed the corresponding full-length cDNA (BOCL3) to be cloned. The deduced amino acid sequence of BOCL3 showed highest homology to the deduced amino acid sequences of several Arabidopsis thaliana genes annotated as tyrosine aminotransferase-like, including a coronatine, jasmonic acid, and salt stress-inducible gene, CORI3 (78.8% identity), and the unidentified rooty/superroot1 gene (44.8% identity). A full-length expressed sequence tag clone of CORI3 was obtained and recombinant CORI3 was synthesized in Escherichia coli. Isolated recombinant CORI3 catalyzed a cystine lyase reaction, but no aminotransferase reactions. The present study identifies, for the first time, a cystine lyase from plants at a molecular level and redefines the functional assignment of the only functionally identified member of a group of A. thaliana genes annotated as tyrosine aminotransferase-like.

Amino Acid Sequence↗

GOAnno: GO annotation based on multiple alignment.

UNLABELLED: GOAnno is a web tool that automatically annotates proteins according to the Gene Ontology (GO) using evolutionary information available in hierarchized multiple alignments. GO terms present in the aligned functional subfamily can be cross-validated and propagated to obtain highly reliable predicted GO annotation based on the GOAnno algorithm. AVAILABILITY: The web tool and a reduced version for local installation are freely available at http://igbmc.u-strasbg.fr/GOAnno/GOAnno.html SUPPLEMENTARY INFORMATION: The website supplies a detailed explanation and illustration of the algorithm at http://igbmc.u-strasbg.fr/GOAnno/GOAnnoHelp.html.

Algorithms↗

A comprehensive update of the sequence and structure classification of kinases.

BACKGROUND: A comprehensive update of the classification of all available kinases was carried out. This survey presents a complete global picture of this large functional class of proteins and confirms the soundness of our initial kinase classification scheme. RESULTS: The new survey found the total number of kinase sequences in the protein database has increased more than three-fold (from 17,310 to 59,402), and the number of determined kinase structures increased two-fold (from 359 to 702) in the past three years. However, the framework of the original two-tier classification scheme (in families and fold groups) remains sufficient to describe all available kinases. Overall, the kinase sequences were classified into 25 families of homologous proteins, wherein 22 families (approximately 98.8% of all sequences) for which three-dimensional structures are known fall into 10 fold groups. These fold groups not only include some of the most widely spread proteins folds, such as the Rossmann-like fold, ferredoxin-like fold, TIM-barrel fold, and antiparallel beta-barrel fold, but also all major classes (all alpha, all beta, alpha+beta, alpha/beta) of protein structures. Fold predictions are made for remaining kinase families without a close homolog with solved structure. We also highlight two novel kinase structural folds, riboflavin kinase and dihydroxyacetone kinase, which have recently been characterized. Two protein families previously annotated as kinases are removed from the classification based on new experimental data. CONCLUSION: Structural annotations of all kinase families are now revealed, including fold descriptions for all globular kinases, making this the first large functional class of proteins with a comprehensive structural annotation. Potential uses for this classification include deduction of protein function, structural fold, or enzymatic mechanism of poorly studied or newly discovered kinases based on proteins in the same family.

Algorithms↗

RNAi-induced phenotypes suggest a novel role for a chemosensory protein CSP5 in the development of embryonic integument in the honeybee (Apis mellifera).

Small chemosensory proteins (CSPs) belong to a conserved, but poorly understood, protein family found in insects and other arthropods. They exhibit both broad and restricted expression patterns during development. In this paper, we used a combination of genome annotation, transcriptional profiling and RNA interference to unravel the functional significance of a honeybee gene (csp5) belonging to the CSP family. We show that csp5 expression resembles the maternal-zygotic pattern that is characterized by the initiation of transcription in the ovary and the replacement of maternal mRNA with embryonic mRNA. Blocking the embryonic expression of csp5 with double-stranded RNA causes abnormalities in all body parts where csp5 is highly expressed. The treated embryos show a "diffuse", often grotesque morphology, and the head skeleton appears to be severely affected. They are 'unable-to-hatch' and cannot progress to the larval stages. Our findings reveal a novel, essential role for this gene family and suggest that csp5 (unable-to-hatch) is an ectodermal gene involved in embryonic integument formation. Our study confirms the utility of an RNAi approach to functional characterization of novel developmental genes uncovered by the honeybee genome project and provides a starting point for further studies on embryonic integument formation in this insect.

Amino Acid Sequence↗

Temporal evolution of the Arabidopsis oxidative stress response.

We have carried out a detailed analysis of the changes in gene expression levels in Arabidopsis thaliana ecotype Columbia (Col-0) plants during and for 6 h after exposure to ozone (O3) at 350 parts per billion (ppb) for 6 h. This O3 exposure is sufficient to induce a marked transcriptional response and an oxidative burst, but not to cause substantial tissue damage in Col-0 wild-type plants and is within the range encountered in some major metropolitan areas. We have developed analytical and visualization tools to automate the identification of expression profile groups with common gene ontology (GO) annotations based on the sub-cellular localization and function of the proteins encoded by the genes, as well as to automate promoter analysis for such gene groups. We describe application of these methods to identify stress-induced genes whose transcript abundance is likely to be controlled by common regulatory mechanisms and summarized our findings in a temporal model of the stress response.

Arabidopsis↗

Sequencing and analysis of the large virulence plasmid pLVPK of Klebsiella pneumoniae CG43.

We have determined the entire DNA sequence of pLVPK, which is a 219-kb virulence plasmid harbored in a bacteremic isolate of Klebsiella pneumoniae. A total of 251 open reading frames (ORFs) were annotated, of which 37% have homologous genes of known function, 31% match the hypothetical genes in the GenBank database, and the remaining 32% are novel sequences. The obvious virulence-associated genes carried by the plasmid are the capsular polysaccharide synthesis regulator rmpA and its homolog rmpA2, and multiple iron-acquisition systems, including iucABCDiutA and iroBCDN siderophore gene clusters, Mesorhizobium loti fepBC ABC-type transporter, and Escherichia coli fecIRA, which encodes a Fur-dependent regulatory system for iron uptake. In addition, several gene clusters homologous with copper, silver, lead, and tellurite resistance genes of other bacteria were also identified. Identification of a replication origin consisting of a repA gene lying in between two sets of iterons suggests that the replication of pLVPK is iteron-controlled and the iterons are the binding sites for the repA to initiate replication and maintain copy number of the plasmid. Genes homologous with E. coli sopA/sopB and parA/parB with nearby direct DNA repeats were also identified indicating the presence of an F plasmid-like partitioning system. Finally, the presence of 13 insertion sequences located mostly at the boundaries of the aforementioned gene clusters suggests that pLVPK was derived from a sequential assembly of various horizontally acquired DNA fragments.

Amino Acid Sequence↗

A survey of expressed tRNA genes in the chromosome I of Arabidopsis using an RNA polymerase III-dependent in vitro transcription system.

Eukaryotic tRNA genes are transcribed by RNA polymerase III. These tRNA genes are generally predicted using computer programs, and 620 tRNA genes in the Arabidopsis thaliana genome are currently annotated. However, no effort has been made to assay whether these predicted tRNA genes are all expressed, because it has been difficult to assay by routine in vivo methods. We report here a large-scale tRNA expression assay of predicted Arabidopsis tRNA genes using an RNA polymerase III-dependent in vitro transcription system developed by our group. DNA fragments including an annotated tRNA gene each were amplified by PCR and the resulting linear DNA was subjected to in vitro transcription. The addition of poly(dA-dT).poly(dA-dT) enhanced activity significantly and reduced background. The 124 predicted tRNA genes present in the Arabidopsis chromosome I were examined, and transcription activity and transcript stability from individual genes were determined. These results indicated that eight annotated genes are not expressed. Based on previous reports on pseudo-tRNA genes (e.g., Beier and Beier, Mol. Gen. Genet. 1992; 233: 201-208) and the present results, we estimated that 16% or more of the annotated tRNA genes in the chromosome I are not functional.

5' Flanking Region↗

Silencing the transcriptome's dark matter: mechanisms for suppressing translation of intergenic transcripts.

Large portions of the genomes of higher eukaryotes are transcribed into RNA molecules that are never destined for translation into proteins. Although some of these transcripts have clearly defined biological roles other than protein coding, most arise from genomic regions devoid of functional genes and many are antisense to regions containing annotated genes. A variety of mechanisms exist to prevent adventitious production of proteins from these transcripts, ranging from degradation within the nucleus to translational silencing in the cytosol.

Animals↗

Computational prediction of the effects of non-synonymous single nucleotide polymorphisms in human DNA repair genes.

Non-synonymous single nucleotide polymorphisms (nsSNPs) represent common genetic variation that alters encoded amino acids in proteins. All nsSNPs may potentially affect the structure or function of expressed proteins and could therefore have an impact on complex diseases. In an effort to evaluate the phenotypic effect of all known nsSNPs in human DNA repair genes, we have characterized each polymorphism in terms of different functional properties. The properties are computed based on amino acid characteristics (e.g. residue volume change); position-specific phylogenetic information from multiple sequence alignments and from prediction programs such as SIFT (Sorting Intolerant From Tolerant) and PolyPhen (Polymorphism Phenotyping). We provide a comprehensive, updated list of all validated nsSNPs from dbSNP (public database of human single nucleotide polymorphisms at National Center for Biotechnology Information, USA) located in human DNA repair genes. The list includes repair enzymes, genes associated with response to DNA damage as well as genes implicated with genetic instability or sensitivity to DNA damaging agents. Out of a total of 152 genes involved in DNA repair, 95 had validated nsSNPs in them. The fraction of nsSNPs that had high probability of being functionally significant was predicted to be 29.6% and 30.9%, by SIFT and PolyPhen respectively. The resulting list of annotated nsSNPs is available online (http://dna.uio.no/repairSNP), and is an ongoing project that will continue assessing the function of coding SNPs in human DNA repair genes.

Computational Biology↗

Complete nucleotide sequence of pSCV50, the virulence plasmid of Salmonella enterica serovar Choleraesuis SC-B67.

We carried out comparative analysis on the sequences of two 50-kb virulence plasmids of Salmonella enterica serovar Choleraesuis strains SC-B67 (pSCV50) and RF-1 (pKDSC50). The two plasmids share over 99% sequence similarity. Ninety-two nucleotide variations at 42 sites were detected between the two plasmids; pSCV50 contains 24 nucleotide substitutions, 6 deletions, and 62 insertions, compared to pKDSC50. Two regions in pSCV50 appeared to be more susceptible to changes: one is the non-virulence-associated transfer region (27.5-33.0 K) and the other a function-unknown region (9.0-10.5 K). We re-annotated pSCV50 using more advanced tools and the up-to-date databases and corrected the inaccurate annotation in pKDSC50. The results indicate that virulence-related genes on the 50-kb plasmid are under negative selection, suggesting that they play important roles in the expression of virulence during the process of infection, while other genes in this plasmid tend to evolve neutrally.

Chromosome Mapping↗

Modeling comparative mapping using objects and associations.

Spatial information on genome organization is essential for both gene prediction and annotation among species and a better understanding of genomes functioning and evolution. We propose in this article an object-association model to formalize comparative genomic mapping. This model is being implemented in the GeMCore knowledge base, for which some original capabilities are described. GeMCore associated to the GeMME graphical interface for molecular evolution was used to spatially characterize the minor shift phenomenon between human and mouse.

Animals↗

The ABC of ABCS: a phylogenetic and functional classification of ABC systems in living organisms.

ATP binding cassette (ABC) systems constitute one of the most abundant superfamilies of proteins. They are involved not only in the transport of a wide variety of substances, but also in many cellular processes and in their regulation. In this paper, we made a comparative analysis of the properties of ABC systems and we provide a phylogenetic and functional classification. This analysis will be helpful to accurately annotate ABC systems discovered during the sequencing of the genome of living organisms and to identify the partners of the ABC ATPases.

ATP-Binding Cassette Transporters↗

Oligonucleotide frequency matrices addressed to recognizing functional DNA sites.

MOTIVATION: Recognition of functional sites remains a key event in the course of genomic DNA annotation. It is well known that a number of sites have their own specific oligonucleotide content. This pinpoints the fact that the preference of the site-specific nucleotide combinations at adjacent positions within an analyzed functional site could be informative for this site recognition. Hence, Web-available resources describing the site-specific oligonucleotide content of the functional DNA sites and applying the above approach for site recognition are needed. However, they have been poorly developed up to now. RESULTS: To describe the specific oligonucleotide content of the functional DNA sites, we introduce the oligonucleotide alphabets, out of which the frequency matrix for a given site could be constructed in addition to a traditional nucleotide frequency matrix. Thus, site recognition accuracy increases. This approach was implemented in the activated MATRIX database accumulating oligonucleotide frequency matrices of the functional DNA sites. We have demonstrated that the false-positive error of the functional site recognition decreases if the oligonucleotide frequency matrixes are added to the nucleotide frequency matrixes commonly used. AVAILABILITY: The MATRIX database is available on the Web, http://wwwmgs.bionet.nsc.ru/Dbases/MATRIX/ and the mirror site, http://www.cbil.upenn.edu/mgs/systems/c onsfreq/.

Algorithms↗

The HIB database of annotated UniGene clusters.

SUMMARY: The HumanInfoBase (HIB) is a database of putative human gene transcripts. UniGene clusters are assembled, and the resulting consensus sequences are submitted to the PEDANT software system (Frishman,D., Albermann,K., Hani,J., Heumann,K., Metanomski,A., Zollner,A. and Mewes,H.-W., 2001, Bioinformatics, 17, 44--57) for fully automatic sequence analysis and annotation. Predicted transcripts are classified using a variety of functional and structural categories, and hyperlinks to various databases are provided for additional information. A WWW-based graphical user interface represents the assembly process as well as functionally important sites in the putative transcripts.

Data Collection↗

Proteome analysis based on motif statistics.

MOTIVATION: Even for the amino acid motifs collected in the Prosite database there may be chance occurences as opposed to those occurences where the motif is involved in fold or function of a protein. With recent mathematical advances in assessing the significance of observing such a motif a particular number of times, we can now study the over- or under-representation of particular motifs in a complete genome and attempt to make functional deductions. RESULTS: We demonstrate that statistical over- or under-representation of motifs in complete proteomes may be an indicator of whether, in that organism, we are looking at chance occurrences of the motif or whether the occurrences are sufficiently numerous to suggest a systematic, and thus functionally important occurrence. This has important implications on databank annotations. AVAILABILITY: The complete dataset comprising the plotted statistics of 266 Prosite motifs on 42 proteomes is available at http://algo.inria.fr/nicodeme/proteomes/proteocomp.html. The software used to compute this data has been described by Nicodème (2000, 2001). They are available either by web access as mentioned in these articles or by direct request from Pierre Nicodème.

Amino Acid Motifs↗

SVbyEye: a visual tool to characterize structural variation among whole-genome assemblies.

MOTIVATION: We are now in the era of being able to routinely generate highly contiguous (near telomere-to-telomere) genome assemblies of human and nonhuman species. Complex structural variation and regions of rapid evolutionary turnover are being discovered for the first time. Thus, efficient and informative visualization tools are needed to evaluate and directly observe structural differences between two or more genomes. RESULTS: We developed SVbyEye, an open-source R package to visualize and annotate sequence-to-sequence alignments along with various functionalities to process these alignments. The tool facilitates the characterization of complex structural variants in the context of sequence homology helping resolve the mechanisms underlying their formation. AVAILABILITY AND IMPLEMENTATION: SVbyEye is available on GitHub (https://github.com/daewoooo/SVbyEye) and via Zenodo (https://doi.org/10.5281/zenodo.15303553).

Software↗

Nallo: a Nextflow pipeline for comprehensive human long-read genome analysis.

MOTIVATION: Long-read sequencing (LRS) is increasingly used for human medical research and clinical diagnostics due to its capacity to generate complete genome information. However, there is a lack of robust and easy-to-use pipelines for comprehensive LRS data analysis. RESULTS: Here we present Nallo, a Nextflow pipeline for analysis of PacBio and Oxford Nanopore data, with additional support for rare disease research projects. The pipeline detects a wide range of genetic variants, performs genome assembly, and reports CpG methylation. It also enables annotation and ranking of variants based on their predicted functional consequences. AVAILABILITY AND IMPLEMENTATION: Nallo is available from GitHub: https://github.com/genomic-medicine-sweden/nallo.

Humans↗

The SBASE protein domain library, release 2.0: a collection of annotated protein sequence segments.

SBASE 2.0 is the second release of SBASE, a collection of annotated protein domain sequences. SBASE entries represent various structural, functional, ligand-binding and topogenic segments of proteins [Pongor, S. et al. (1993) Prot. Eng., in press]. This release contains 34,518 entries provided with standardized names and it is cross-referenced to the major protein and nucleic acid databanks as well as to the PROSITE catalog of protein sequence patterns [Bairoch, A. (1992) Nucl. Acids Res., 20 suppl, 2013-2018]. SBASE can be used for establishing domain homologies using different database-search tools such as FASTA [Lipman and Pearson (1985) Science, 227, 1436-1441], FASTDB [Brutlag et al. (1990) Comp. Appl. Biosci., 6, 237-245] or BLAST3 [Altschul and Lipman (1990) Proc. Natl. Acad. Sci. USA, 87, 5509-5513] which is especially useful in the case of loosely defined domain types for which efficient consensus patterns can not be established. SBASE 2.0 and a set of search and retrieval tools are freely available on request to the authors or by anonymous 'ftp' file transfer from mean value of ftp.icgeb.trieste.it.

Amino Acid Sequence↗