Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 991 records · Page 55Linked to original sources

Identification of novel clock-controlled genes by cDNA macroarray analysis in Chlamydomonas reinhardtii.

Circadian rhythms are self-sustaining oscillations whose period length under constant conditions is about 24 h. Circadian rhythms are widespread and involve functions as diverse as human sleep-wake cycles and cyanobacterial nitrogen fixation. In spite of a long research history, knowledge about clock-controlled genes is limited in Chlamydomonas reinhardtii. Using a cDNA macroarray containing 10 368 nuclear-encoded genes, we examined global circadian regulation of transcription in Chlamydomonas. We identified 269 candidates for circadianly expressed gene. Northern blot analysis confirmed reproducible and sustainable rhythmicity for 12 genes. Most genes exhibited peak expression at the transition point between day and night. One hundred and eighteen genes were assigned predicted annotations. The functions of the cycling genes were diverse and included photosynthesis, respiration, cellular structure, and various metabolic pathways. Surprisingly, 18 genes encoding chloroplast ribosomal proteins showed a coordinated circadian pattern of expression and peaked just at the beginning of subjective day. The co-regulation of genes bearing a similar function was also observed in genes involved in cellular structure. They peaked at the end of the subjective night, which is when the regeneration of cell walls and flagella in daughter cells occurs. Expression of the chlamyopsin gene, which encodes an opsin-type photoreceptor, also exhibited circadian rhythm.

Animals↗

Chromosome-level genome assembly and annotation of Pterygoplichthys pardalis.

Suckermouth catfishes, with their evolved powerful features, have become notorious invasive species, causing significant damage to aquatic ecosystems. However, the lack of high-quality genomes severely restricts research on this group within the field. In this study, we de novo assembled the chromosome-level genome assembly of Pterygoplichthys pardalis using multiple platforms of sequencing data, including Illumina short reads, Nanopore long reads, and Hi-C sequencing reads, resulting in a 1.51 Gb genome assembly. Multiple evaluations, including read mapping ratio (98.52%), transcript mapping ratio (99.61%), conserved BUSCO gene set (98.8%), and N50 score (49.47 Mb), indicated the high continuity and accuracy of the genome assembly we generated. Genome annotation found that 0.97 Gb of genome sequences are repetitive sequences, accounting for 64.47% of the genome assembly. Further, 23,859 protein-coding genes were successfully predicted, 92.92% of which could be annotated in functional databases. This high-quality genome assembly of P. pardalis provides a valuable resource for understanding the genetic underpinnings of P. pardalis's invasive success and offers critical data for future fisheries research and management.

Animals↗

Genomic characterization of a repetitive motif strongly associated with developmental genes in Drosophila.

BACKGROUND: Non-coding DNA represents a high proportion of all metazoan genomes. Although an undetermined fraction of this DNA may be considered devoid of any function, it also contains important information residing in specific cis-regulatory sequences. RESULTS: We report a 27 bp motif that is overrepresented within the fly genome. This motif does not show any significant similarity with transposon sequences and is strongly associated with genes involved in development and/or signal transduction. The 27 bp motif is preferentially located within introns, and has a tendency to be present in multiple copies around genes. Furthermore, it is often found embedded in known non-coding regulatory regions. The regulatory network defined by this motif is partially shared in D. pseudoobscura. CONCLUSION: We have identified a 27 bp cis-regulatory sequence widely distributed within the Drosophila genome in association with developmental genes. This motif may be very useful towards the annotation of functional regulatory regions within the Drosophila genome and the construction of regulatory networks of Drosophila development.

Animals↗

Annotation and expression profile analysis of 2073 full-length cDNAs from stress-induced maize (Zea mays L.) seedlings.

Full-length cDNAs are very important for genome annotation and functional analysis of genes. The number of full-length cDNAs from maize (Zea mays L.) remains limited. Here we report the construction of a full-length enriched cDNA library from osmotically stressed maize seedlings by using the modified CAP trapper method. From this library, 2073 full-length cDNAs were collected and further analyzed by sequencing from both the 5'- and 3'-ends. A total of 1728 (83.4%) sequences did not match known maize mRNA and full-length cDNA sequences in the GenBank database and represent new full-length genes. After alignment of the 2073 full-length cDNAs with 448 maize BAC sequences, it was found that 84 full-length cDNAs could be mapped to the BACs. Of these, 43 genes (51.2%) have been correctly annotated from the BAC clones, 37 genes (44.0%) have been annotated with a different exon-intron structure from our cDNA, and four genes (4.76%) had no annotations in the TIGR database. Expression analysis of 2073 full-length maize cDNAs using a cDNA macroarray led to the identification of 79 genes upregulated by stress treatments and 329 downregulated genes. Of the 79 stress-inducible genes, 30 genes contain ABRE, DRE, MYB, MYC core sequences or other abiotic-responsive cis-acting elements in their promoters. These results suggest that these cis-acting elements and the corresponding transcription factors take part in plant responses to osmotic stress either cooperatively or independently. Additionally, the data suggest that an ethylene signaling pathway may be involved in the maize response to drought stress.

DNA, Complementary↗

Duplicated genes evolve slower than singletons despite the initial rate increase.

BACKGROUND: Gene duplication is an important mechanism that can lead to the emergence of new functions during evolution. The impact of duplication on the mode of gene evolution has been the subject of several theoretical and empirical comparative-genomic studies. It has been shown that, shortly after the duplication, genes seem to experience a considerable relaxation of purifying selection. RESULTS: Here we demonstrate two opposite effects of gene duplication on evolutionary rates. Sequence comparisons between paralogs show that, in accord with previous observations, a substantial acceleration in the evolution of paralogs occurs after duplication, presumably due to relaxation of purifying selection. The effect of gene duplication on evolutionary rate was also assessed by sequence comparison between orthologs that have paralogs (duplicates) and those that do not (singletons). It is shown that, in eukaryotes, duplicates, on average, evolve significantly slower than singletons. Eukaryotic ortholog evolutionary rates for duplicates are also negatively correlated with the number of paralogs per gene and the strength of selection between paralogs. A tally of annotated gene functions shows that duplicates tend to be enriched for proteins with known functions, particularly those involved in signaling and related cellular processes; by contrast, singletons include an over-abundance of poorly characterized proteins. CONCLUSIONS: These results suggest that whether or not a gene duplicate is retained by selection depends critically on the pre-existing functional utility of the protein encoded by the ancestral singleton. Duplicates of genes of a higher biological import, which are subject to strong functional constraints on the sequence, are retained relatively more often. Thus, the evolutionary trajectory of duplicated genes appears to be determined by two opposing trends, namely, the post-duplication rate acceleration and the generally slow evolutionary rate owing to the high level of functional constraints.

Animals↗

GXXXG and GXXXA motifs stabilize FAD and NAD(P)-binding Rossmann folds through C(alpha)-H... O hydrogen bonds and van der waals interactions.

Here we present evidence that domains in soluble proteins containing either the GXXXG or GXXXA motif are stabilized by the interaction of a beta-strand with the following alpha-helix. As an example, we characterized a beta-strand-helix interaction from the FAD or NAD(P)-binding Rossmann fold. The Rossmann fold is one of the three most highly represented folds in the Protein Data Bank (PDB). A subset of the proteins that adopt the Rossmann fold also bind to nucleotide cofactors such as FAD and NAD(P) and function as oxidoreductases. These Rossmann folds can often be identified by the short amino acid sequence motif, GX(1-2)GXXG. Here, we present evidence that in addition to this sequence motif, Rossmann folds that bind FAD and NAD(P) also typically contain either GXXXG or GXXXA motifs, where the first glycyl residue of these motifs and the third glycyl residue of the GX(1-2)GXXG motif are the same residue. These two motifs appear to stabilize the Rossmann fold: the first glycyl residue of either the GXXXG or GXXXA motif contacts the carbonyl oxygen atom from the first glycyl residue of the GX(1-2)GXXG motif consistent with the formation of a C(alpha)-H cdots, three dots, centered O hydrogen bond. In addition, both the glycyl and alanyl residues of the GXXXG or GXXXA motifs form van der Waals interactions with either a valine or isoleucine residue located either seven or eight residues further back along the polypeptide chain from the first glycine of the GXXXG or GXXXA motifs. Therefore, we combine both the GX(1-2)GXXG and GXXXG/A motifs into an extended motif, V/IXGX(1-2)GXXGXXXG/A, that is more strongly indicative than previously described motifs of Rossmann folds that bind FAD or NAD(P). The V/IXGX(1-2)GXXGXXXG/A motif can be used to search genomic sequence data and to annotate the function of proteins containing the motif as oxidoreductases, including proteins of previously unknown function.

Amino Acid Motifs↗

PICRUSt2-SC: an update to the reference database used for functional prediction within PICRUSt2.

SUMMARY: PICRUSt2 is a bioinformatic tool that predicts microbial functions in amplicon sequencing data using a database of annotated reference genomes. We have constructed an updated database for PICRUSt2 that has substantially increased the number of bacterial (19,493 to 26,868) and archaeal (406 to 1,002) genomes as well as the number of functional annotations present. The previous PICRUSt2 database relied on many timely and computationally intensive manual processes that made it difficult to update. We constructed a new streamlined process to allow regular upgrades to the PICRUSt2 database on an ongoing basis, and used this process to create a new database, PICRUSt2-SC (Sugar-Coated). Additionally, we have shown that this updated database contains genomes that more closely match study sequences from a range of different environments. The genomes contained in the database therefore better represent these environments and this leads to an improvement in the predicted functional annotations obtained from PICRUSt2. AVAILABILITY AND IMPLEMENTATION: PICRUSt2 source code is freely available at https://github.com/picrust/picrust2 and at https://anaconda.org/bioconda/picrust2. The latest version of PICRUSt2 at the time of writing is also archived: https://doi.org/10.5281/zenodo.15119781. The PICRUSt2-SC database comes pre-installed with PICRUSt2 from version 2.6.0 onwards. Step-by-step instructions for making the updated database are at https://github.com/picrust/picrust2/wiki/Updating-the-PICRUSt2-database. All code used for the analyses and figures in this manuscript is at https://github.com/R-Wright-1/PICRUSt2-SC_application_note and https://doi.org/10.5281/zenodo.15119770.

Software↗

Recent innovations in tissue-specific gene modifications in the mouse.

Annotating the functions of individual genes in in vivo contexts has become the primary task of mouse genetics in the post-genome era. In addition to conventional approaches using transgenic technologies and gene targeting, the recent development of conditional gene modification techniques has opened novel opportunities for elucidating gene function at the level of the whole mouse to individual tissues or cell types. Tissue-specific gene modifications in the mouse have been made possible using site-specific DNA recombinases and conditional alleles. Recent innovations in this basic technology have facilitated new types of experiments, revealing novel insights into mammalian embryology. In this review, we focus on these recent innovations and new technical issues that impact the success of these conditional gene modification approaches.

Animals↗

Gene targeting of ErbB3 using a Cre-mediated unidirectional DNA inversion strategy.

Recombinase-mediated unidirectional DNA inversion and transcriptional arrest is a promising strategy for high throughput conditional mutagenesis in the mouse. Banks of mouse embryonic stem cells with defined, transcriptionally silent insertions that can be activated by Cre recombinase would take advantage of existing transgenic Cre lines to rapidly produce hundreds of lineage specific and temporally controlled knockout mice for each gene, thereby introducing significant parallelism to functional gene annotation. However, the extent to which this strategy results in effective gene knockout has not been established. To test the feasibility of this strategy we targeted ErbB3, a member of the ErbB family of tyrosine kinase receptors, using this strategy. Insertion of a reversed "flipflox" vector consisting of a gene inactivation cassette (GI) and an internal ribosome entry site (IRES)-GFP reporter into intron 1 of ErbB3 was transcriptionally silent and did not affect ErbB3 expression. Crosses with ubiquitous and lineage specific Cre recombinase expressing lines permanently inverted the inserted GI cassette and blocked ErbB3 expression. Unidirectional DNA inversion by in vivo recombination is an effective strategy for targeted or ubiquitous gene knockout.

Animals↗

Identifying cysteines and histidines in transition-metal-binding sites using support vector machines and neural networks.

Accurate predictions of metal-binding sites in proteins by using sequence as the only source of information can significantly help in the prediction of protein structure and function, genome annotation, and in the experimental determination of protein structure. Here, we introduce a method for identifying histidines and cysteines that participate in binding of several transition metals and iron complexes. The method predicts histidines as being in either of two states (free or metal bound) and cysteines in either of three states (free, metal bound, or in disulfide bridges). The method uses only sequence information by utilizing position-specific evolutionary profiles as well as more global descriptors such as protein length and amino acid composition. Our solution is based on a two-stage machine-learning approach. The first stage consists of a support vector machine trained to locally classify the binding state of single histidines and cysteines. The second stage consists of a bidirectional recurrent neural network trained to refine local predictions by taking into account dependencies among residues within the same protein. A simple finite state automaton is employed as a postprocessing in the second stage in order to enforce an even number of disulfide-bonded cysteines. We predict histidines and cysteines in transition-metal-binding sites at 73% precision and 61% recall. We observe significant differences in performance depending on the ligand (histidine or cysteine) and on the metal bound. We also predict cysteines participating in disulfide bridges at 86% precision and 87% recall. Results are compared to those that would be obtained by using expert information as represented by PROSITE motifs and, for disulfide bonds, to state-of-the-art methods.

Amino Acid Sequence↗

Assessment of genomic prediction capabilities of transcriptome data in a barley multi-parent RIL population.

Low-cost and high-throughput RNA sequencing data for barley RILs achieved GP performance comparable to or better than traditional SNP array datasets when combined with parental whole-genome sequencing SNP data. The field of genomic selection (GS) is advancing rapidly on many fronts including the utilization of multi-omics datasets with the goal of increasing prediction ability and becoming an integral part of an increasing number of breeding programs ensuring future food security. In this study, we used RNA sequencing (RNA-Seq) data to perform genomic prediction (GP) on three related barley RIL populations. We investigated the potential of increasing prediction ability by combining genomic and transcriptomic datasets, adding whole-genome sequencing (WGS) SNP data, functional annotation-based filtering, and empirical quality filtering. Our RNA-Seq data were generated cost-efficiently using small-footprint plant cultivation, high-throughput RNA extraction, and Library preparation miniaturization. We also examined sequencing depth reduction as an additional cost-saving measure. We used fivefold cross-validation to evaluate the prediction ability of the gene expression dataset, the RNA-Seq SNP dataset, and the consensus SNP dataset between the RNA-Seq and parental WGS data, resulting in prediction abilities between 0.73 and 0.78. The consensus SNP dataset performed best, with five out of eight traits performing significantly better compared to a 50K SNP array, which served as a benchmark. The advantage of the consensus SNP dataset was most prominent in the inter-population predictions, in which the training and validation sets originated from different RIL sub-populations. We were therefore able to not only show that RNA-Seq data alone are able to predict various complex traits in barley using RILs, but also that the performance can be further increased with WGS data for which the public availability will steadily increase.

Hordeum↗

A structure-based method for identifying DNA-binding proteins and their sites of DNA-interaction.

A classification model of a DNA-binding protein chain was created based on identification of alpha helices within the chain likely to bind to DNA. Using the model, all chains in the Protein Data Bank were classified. For many of the chains classified with high confidence, previous documentation for DNA-binding was found, yet no sequence homology to the structures used to train the model was detected. The result indicates that the chain model can be used to supplement sequence based methods for annotating the function of DNA-binding. Four new candidates for DNA-binding were found, including two structures solved through structural genomics efforts. For each of the candidate structures, possible sites of DNA-binding are indicated by listing the residue ranges of alpha helices likely to interact with DNA.

Binding Sites↗

The Paris-Sud yeast structural genomics pilot-project: from structure to function.

We present here the outlines and results from our yeast structural genomics (YSG) pilot-project. A lab-scale platform for the systematic production and structure determination is presented. In order to validate this approach, 250 non-membrane proteins of unknown structure were targeted. Strategies and final statistics are evaluated. We finally discuss the opportunity of structural genomics programs to contribute to functional biochemical annotation.

Genomics↗

Computational identification and sequence analysis of stop codon readthrough genes in Oryza sativa.

Using an approach based on the Readthrough Candidate Extraction System (RCES), we extracted 111 candidates from 9620 gene sequences of rice. The results of homology search and sequence analysis demonstrated that these candidates included actual readthrough genes that would be important for further investigating the mechanism of translation termination regulated by readthrough event, and could also give some useful clues for functional genome annotation. Between the candidates and non-candidates of gene sequences in rice, there exist significant base biases at the positions surrounding the stop codons. These positions, especially both -1 and +4, are referred to as part of an extended stop signal. In candidates, G at position -1, and G or C at position +4 are much more favored than that in non-candidates. Both stop sequence patterns, GUAGC and GUGAG, might drive high readthrough efficiency in rice. Secondary structure analysis revealed that the -1 and +1 amino acids around the first stop codon of candidates have a strong bias toward arginine, particularly the +1 position (20.7%), which indicated that the amino acids at the readthrough region being frequently located in the hydrophilic region of beta-turn might be a determinant for efficient translation termination or not.

Base Sequence↗

A multi-model genome-wide association study identifies genetic variants underlying resistance to Largemouth Bass Ranavirus (LMBV) in Micropterus salmoides.

Largemouth bass (Micropterus salmoides) is an economically important freshwater aquaculture species, yet recurrent outbreaks of Largemouth Bass Ranavirus (LMBV) continue to impair production and cause substantial losses. The genetic basis of host variation in LMBV resistance remains insufficiently characterized. Here, we applied a multi-model genome-wide association study (GWAS) to identify loci associated with resistance following a controlled challenge with the LMBV-23PY strain. Whole-genome resequencing was performed for 146 phenotyped fish, including 72 susceptible and 74 resistant individuals. After stringent quality control, 877,262 high-quality variants were retained and tested using six GWAS models. Across binary survival status and survival time phenotypes, 32 shared suggestive variants were consistently detected across models, representing suggestive loci for LMBV-23PY resistance. Genes within ±50 kb of these loci were annotated, and functional enrichment highlighted immune- and redox-related biological processes. Three prioritized candidates-GSTT3L (glutathione S-transferase theta-3-like), CGRP2 (calcitonin gene-related peptide 2), and NPPC (natriuretic peptide C)-were associated with pathways involved in oxidative stress responses and immune regulation. Collectively, these results provide insight into the genetic architecture of LMBV-23PY resistance in largemouth bass and identify suggestive variants and associated candidate genes for downstream validation, functional interrogation, and the development of marker-assisted and genome-enabled breeding strategies.

Animals↗

SMAUG is a major regulator of maternal mRNA destabilization in Drosophila and its translation is activated by the PAN GU kinase.

In animals, egg activation triggers a cascade of posttranscriptional events that act on maternally synthesized RNAs. We show that, in Drosophila, the PAN GU (PNG) kinase sits near the top of this cascade, triggering translation of SMAUG (SMG), a multifunctional posttranscriptional regulator conserved from yeast to humans. Although PNG is required for cytoplasmic polyadenylation of smg mRNA, it regulates translation via mechanisms that are independent of its effects on the poly(A) tail. Analyses of mutants suggest that PNG relieves translational repression by PUMILIO (PUM) and one or more additional factors, which act in parallel through the smg mRNA's 3' untranslated region (UTR). Microarray-based gene expression profiling shows that SMG is a major regulator of maternal transcript destabilization. SMG-dependent mRNAs are enriched for gene ontology annotations for function in the cell cycle, suggesting a possible causal relationship between failure to eliminate these transcripts and the cell cycle defects in smg mutants.

3' Untranslated Regions↗

Computational tools for the analysis of heteroatom groups and their neighbours in protein tertiary structure.

A number of Protein Data Bank (PDB) entries contain heteroatoms defined as HETATM. These include the atomic co-ordinates mainly for heteroatom groups, such as cofactors, coenzymes, prosthetic groups, metal ions, sugars, drugs, peptides, heavy-atom derivatives, non-standard amino acid residues/nucleotides, water molecules and so on. In order to evaluate the different heteroatom (Het) groups and their distribution in protein tertiary structure, we have extracted these from all proteins in the PDB and provided the data in an easily accessible format at the following website. The data can be queried on the PDB code, protein name/description, Het Group code or Het Group name. Further, we have also developed a web-based software application that reports neighbouring atoms evaluated by a "user-defined" distance cut-off value (in Angstrom units), either between a specific Het Group or all Het Groups in a given PDB with amino acid residues and water molecules in the corresponding protein, or neighbours for only all the amino acid residues in the given PDB with respect to Het Groups and water molecules. Together, the database and software applications are useful to gather information that can be further analyzed in order to obtain insights into the preferred interactions of heteroatom groups in proteins, study their binding mode, design novel molecules or to annotate protein function.

Animals↗

Cholesterol Metabolism-related Characteristics Predict Therapeutic Response and Survival in Esophageal Cancer.

INTRODUCTION: Cholesterol homeostasis has been identified as an essential downstream pathway of mutations in TP53. Esophageal cancer is one of the most prevalent malignancies exhibiting the mutation. OBJECTIVES: To explore the significance of cholesterol metabolism-related characteristics in tumor phenotype and treatment outcomes of esophageal cancer. METHODS: We established a cholesterol metabolism-related gene set (CMGs) and performed Lasso-Cox analysis to identify prognostic signatures. Nomogram-based risk scores and clinical stages afterwards were constructed and evaluated. We simultaneously identified two metabolic subtypes based on the distinct features of the CMGs. We annotated the functional and pathway characteristics of differentially expressed genes between the clusters and compared the differences in clinical and immune characteristics. Finally, we assessed the prognostic value of signatures in the GSE53625 and two clinical cohorts using whole-exon sequencing and multiplex immunofluorescence. RESULTS: Our study identified five cholesterol prognosis-related genes (CRGs) that demonstrated superior prognostic efficacy in the training set compared to clinical staging, validated in independent public databases and two clinical cohorts. According to the different expression patterns of the signatures, patients were divided into two subtypes. The C1 group demonstrated poorer overall survival, response to immunotherapy, and downregulation of the p53 pathway. In the immune correlation analysis, we found that the risk score based on 5-signature model was significantly positively correlated with the abundance of suppressive immune cells and the immune checkpoints. Finally, we explored the impact of expression and genomic polymorphism of the signatures on the prognosis at the pan-cancer level. CONCLUSIONS: Our findings underscore the distinct expression patterns of CRGs in esophageal cancer. These signatures are efficient to serve as prognostic indicators and assess the effectiveness of immunotherapy. They may also represent promising targets in other TP53 mutant malignancies.

Humans↗