Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Serum Proteomic Profiling Reveals Renin-Associated Immune and Cytoskeletal Dysregulation in Post-COVID-19 Condition Patients with Secondary Adrenal Insufficiency.

Post-COVID-19 condition (PCC) with secondary adrenal insufficiency (SAI) involves multiorgan dysfunction, potentially linked to renin-angiotensin-aldosterone system dysregulation. The molecular basis of renin-associated pathology remains unclear. Here, PCC+SAI patients were stratified by upright renin into low- (<38.8&#x202f;pg/mL) and high-renin (&#x2265;38.8&#x202f;pg/mL) groups. Clinical, endocrine, and proteomic analyses were performed. We found that high-renin patients showed increased BMI, lipids, renin, and aldosterone, but reduced aldosterone-to-renin ratio. Proteomic annalysis identified 20 differentially expressed proteins (DEPs), including 17 upregulated and 3 downregulated proteins in Ren-H patients. Functional annotation revealed that 15 DEPs were immune-related (e.g., APOC4, APOE, C4BPA, CFAH, CFHR3, PF4V, PLF4), while FLNA and COF1 represented cytoskeletal proteins. These DEPs were primarily involved in immune response, complement and coagulation cascades, and MAPK signaling pathways. Correlation analyses indicated that upright renin was positively correlated with complement-related proteins and platelet-derived immune factors, while cytoskeletal proteins (FLNA, COF1) showed positive associations with serum Na+ levels. Additionally, white blood cell and platelet counts were positively correlated with the majority of DEPs. In conclusion, exploratory proteomic analyses suggest that elevated upright renin in PCC+SAI may be associated with immune dysregulation, complement activation, and cytoskeletal remodeling, offering novel insights into the endocrine-immune interactions driving postviral sequelae.

Humans↗

Coupled folding and binding with alpha-helix-forming molecular recognition elements.

Many protein-protein and protein-nucleic acid interactions involve coupled folding and binding of at least one of the partners. Here, we propose a protein structural element or feature that mediates the binding events of initially disordered regions. This element consists of a short region that undergoes coupled binding and folding within a longer region of disorder. We call these features "molecular recognition elements" (MoREs). Examples of MoREs bound to their partners can be found in the alpha-helix, beta-strand, polyproline II helix, or irregular secondary structure conformations, and in various mixtures of the four structural forms. Here we describe an algorithm that identifies regions having propensities to become alpha-helix-forming molecular recognition elements (alpha-MoREs) based on a discriminant function that indicates such regions while giving a low false-positive error rate on a large collection of structured proteins. Application of this algorithm to databases of genomics and functionally annotated proteins indicates that alpha-MoREs are likely to play important roles protein-protein interactions involved in signaling events.

Binding Sites↗

Information decay in molecular docking screens against holo, apo, and modeled conformations of enzymes.

Molecular docking uses the three-dimensional structure of a receptor to screen a small molecule database for potential ligands. The dependence of docking screens on the conformation of the binding site remains an open question. To evaluate the information loss that occurs as the active site conformation becomes less defined, a small molecule database was docked against the holo (ligand bound), apo, and homology modeled structures of 10 different enzyme binding sites. The holo and apo representations were crystallographic structures taken from the Protein Data Bank (PDB), and the homology-modeled structures were taken from the publicly available resource ModBase. The database docked was the MDL Drug Data Report (MDDR), a functionally annotated database of 95000 small molecules that contained at least 35 ligands for each of the 10 systems. In all sites, at least 99% of the molecules in the MDDR were treated as nonbinding decoys. For each system, the holo, apo, and modeled structures were used to screen the MDDR, and the ability of each structure to enrich the known ligands for that system over random selection was evaluated. The best overall enrichment was produced by the holo structure in seven systems, the apo structure in two systems, and the modeled structure in one system. These results suggest that the performance of the docking calculation is affected by the particular representation of the receptor used in the screen, and that the holo structure is the one most likely to yield the best discrimination between known ligands and decoy molecules, but important exceptions to this rule also emerge from this study. Although each of the holo, apo, and modeled conformations led to enrichment of known ligands in all systems, the enrichment did not always rise to a level judged to be sufficient to justify the effort of a docking screen. Using a 20-fold enrichment of known ligands over random selection as a rough guideline for what might be enough to justify a docking screen, the holo conformation of the enzyme met this criterion in eight of 10 sites, whereas the apo conformation met this criterion in only two sites and the modeled conformation in three.

Animals↗

Phosphoproteome analysis of HeLa cells using stable isotope labeling with amino acids in cell culture (SILAC).

Identification of phosphorylated proteins remains a difficult task despite technological advances in protein purification methods and mass spectrometry. Here, we report identification of tyrosine-phosphorylated proteins by coupling stable isotope labeling with amino acids in cell culture (SILAC) to mass spectrometry. We labeled HeLa cells with stable isotopes of tyrosine, or, a combination of arginine and lysine to identify tyrosine phosphorylated proteins. This allowed identification of 118 proteins, of which only 45 proteins were previously described as tyrosine-phosphorylated proteins. A total of 42 in vivo tyrosine phosphorylation sites were mapped, including 34 novel ones. We validated the phosphorylation status of a subset of novel proteins including cytoskeleton associated protein 1, breast cancer anti-estrogen resistance 3, chromosome 3 open reading frame 6, WW binding protein 2, Nice-4 and RNA binding motif protein 4. Our strategy can be used to identify potential kinase substrates without prior knowledge of the signaling pathways and can also be applied to profiling to specific kinases in cells. Because of its sensitivity and general applicability, our approach will be useful for investigating signaling pathways in a global fashion and for using phosphoproteomics for functional annotation of genomes.

Amino Acid Sequence↗

Detecting differential and correlated protein expression in label-free shotgun proteomics.

Recent studies have revealed a relationship between protein abundance and sampling statistics, such as sequence coverage, peptide count, and spectral count, in label-free liquid chromatography-tandem mass spectrometry (LC-MS/MS) shotgun proteomics. The use of sampling statistics offers a promising method of measuring relative protein abundance and detecting differentially expressed or coexpressed proteins. We performed a systematic analysis of various approaches to quantifying differential protein expression in eukaryotic Saccharomyces cerevisiae and prokaryotic Rhodopseudomonas palustris label-free LC-MS/MS data. First, we showed that, among three sampling statistics, the spectral count has the highest technical reproducibility, followed by the less-reproducible peptide count and relatively nonreproducible sequence coverage. Second, we used spectral count statistics to measure differential protein expression in pairwise experiments using five statistical tests: Fisher's exact test, G-test, AC test, t-test, and LPE test. Given the S. cerevisiae data set with spiked proteins as a benchmark and the false positive rate as a metric, our evaluation suggested that the Fisher's exact test, G-test, and AC test can be used when the number of replications is limited (one or two), whereas the t-test is useful with three or more replicates available. Third, we generalized the G-test to increase the sensitivity of detecting differential protein expression under multiple experimental conditions. Out of 1622 identified R. palustris proteins in the LC-MS/MS experiment, the generalized G-test detected 1119 differentially expressed proteins under six growth conditions. Finally, we studied correlated expression of these 1119 proteins by analyzing pairwise expression correlations and by delineating protein clusters according to expression patterns. Through pairwise expression correlation analysis, we demonstrated that proteins co-located in the same operon were much more strongly coexpressed than those from different operons. Combining cluster analysis with existing protein functional annotations, we identified six protein clusters with known biological significance. In summary, the proposed generalized G-test using spectral count sampling statistics is a viable methodology for robust quantification of relative protein abundance and for sensitive detection of biologically significant differential protein expression under multiple experimental conditions in label-free shotgun proteomics.

Bacterial Proteins↗

Engineering chromosomal rearrangements in mice.

The combination of gene-targeting techniques in mouse embryonic stem cells and the Cre/loxP site-specific recombination system has resulted in the emergence of chromosomal-engineering technology in mice. This advance has opened up new opportunities for modelling human diseases that are associated with chromosomal rearrangements. It has also led to the generation of visibly marked deletions and balancer chromosomes in mice, which provide essential reagents for maximizing the efficiency of large-scale mutagenesis efforts and which will accelerate the functional annotation of mammalian genomes, including the human genome.

Animals↗

Functional organization of the yeast proteome by systematic analysis of protein complexes.

Most cellular processes are carried out by multiprotein complexes. The identification and analysis of their components provides insight into how the ensemble of expressed proteins (proteome) is organized into functional units. We used tandem-affinity purification (TAP) and mass spectrometry in a large-scale approach to characterize multiprotein complexes in Saccharomyces cerevisiae. We processed 1,739 genes, including 1,143 human orthologues of relevance to human biology, and purified 589 protein assemblies. Bioinformatic analysis of these assemblies defined 232 distinct multiprotein complexes and proposed new cellular roles for 344 proteins, including 231 proteins with no previous functional annotation. Comparison of yeast and human complexes showed that conservation across species extends from single proteins to their molecular environment. Our analysis provides an outline of the eukaryotic proteome as a network of protein complexes at a level of organization beyond binary interactions. This higher-order map contains fundamental biological information and offers the context for a more reasoned and informed approach to drug discovery.

Cells, Cultured↗

Silencing of microRNAs in vivo with 'antagomirs'.

MicroRNAs (miRNAs) are an abundant class of non-coding RNAs that are believed to be important in many biological processes through regulation of gene expression. The precise molecular function of miRNAs in mammals is largely unknown and a better understanding will require loss-of-function studies in vivo. Here we show that a novel class of chemically engineered oligonucleotides, termed 'antagomirs', are efficient and specific silencers of endogenous miRNAs in mice. Intravenous administration of antagomirs against miR-16, miR-122, miR-192 and miR-194 resulted in a marked reduction of corresponding miRNA levels in liver, lung, kidney, heart, intestine, fat, skin, bone marrow, muscle, ovaries and adrenals. The silencing of endogenous miRNAs by this novel method is specific, efficient and long-lasting. The biological significance of silencing miRNAs with the use of antagomirs was studied for miR-122, an abundant liver-specific miRNA. Gene expression and bioinformatic analysis of messenger RNA from antagomir-treated animals revealed that the 3' untranslated regions of upregulated genes are strongly enriched in miR-122 recognition motifs, whereas downregulated genes are depleted in these motifs. Analysis of the functional annotation of downregulated genes specifically predicted that cholesterol biosynthesis genes would be affected by miR-122, and plasma cholesterol measurements showed reduced levels in antagomir-122-treated mice. Our findings show that antagomirs are powerful tools to silence specific miRNAs in vivo and may represent a therapeutic strategy for silencing miRNAs in disease.

3' Untranslated Regions↗

Probabilistic model of the human protein-protein interaction network.

A catalog of all human protein-protein interactions would provide scientists with a framework to study protein deregulation in complex diseases such as cancer. Here we demonstrate that a probabilistic analysis integrating model organism interactome data, protein domain data, genome-wide gene expression data and functional annotation data predicts nearly 40,000 protein-protein interactions in humans-a result comparable to those obtained with experimental and computational approaches in model organisms. We validated the accuracy of the predictive model on an independent test set of known interactions and also experimentally confirmed two predicted interactions relevant to human cancer, implicating uncharacterized proteins into definitive pathways. We also applied the human interactome network to cancer genomics data and identified several interaction subnetworks activated in cancer. This integrative analysis provides a comprehensive framework for exploring the human protein interaction network.

Chromosome Mapping↗

Insertional mutagenesis in mice: new perspectives and tools.

Insertional mutagenesis has been at the core of functional genomics in many species. In the mouse, improved vectors and methodologies allow easier genome-wide and phenotype-driven insertional mutagenesis screens. The ability to generate homozygous diploid mutations in mouse embryonic stem cells allows prescreening for specific null phenotypes prior to in vivo analysis. In addition, the discovery of active transposable elements in vertebrates, and their development as genetic tools, has led to in vivo forward insertional mutagenesis screens in the mouse. These new technologies will greatly contribute to the speed and ease with which we achieve complete functional annotation of the mouse genome.

Animals↗

Proinsulin regulators identified with CRISPR screen and in vivo mouse QTL mapping.

Altered proinsulin levels in &#x3b2;-cells and bloodstream are hallmarks of diabetes and other diseases, but our knowledge about the proinsulin regulators remains limited. Here we perform a genome-wide CRISPR screen to identify 84 proinsulin regulators that alter intracellular proinsulin/insulin ratio in a mouse &#x3b2;-cell line. The proinsulin regulators are distinct from the insulin regulators from a previous orthogonal CRISPR screen. Functional annotation of the proinsulin regulators highlights Golgi as the primary organelle for proinsulin storage and regulation. Trafficking towards the Golgi increases the intra-cellular proinsulin/insulin ratio, while trafficking away from the Golgi, including exocytosis and Golgi-to-ER retrograde transport, decreases the intracellular proinsulin levels. We also map mouse quantitative trait loci (QTLs) associated with plasma proinsulin levels and use the CRISPR screen results to pinpoint the causal genes within the QTL loci. Interestingly, protein disulfide isomerase Pdia6 is the strongest hit from both CRISPR screen and the in vivo QTL mapping. Knocking down Pdia6 significantly reduce proinsulin accumulation in Golgi and secretory granules. Intriguingly, Pdia6-depletion in both human and mouse &#x3b2;-cells does not affect the folding status of proinsulin but causes significantly impaired proinsulin production through a UPR-independent mechanism. Taken together, our genetic profiles provide mechanistic insights into the regulation of proinsulin/insulin homeostasis.

Animals↗

Cell-type signatures of Alzheimer's disease shared across population groups.

Genomic studies at single-cell resolution have identified several cell types associated with clinical and pathological traits in Alzheimer's disease1-9, but have not examined associations that are shared across populations. To bridge this gap, here we use single-nucleus RNA sequencing and assay for transposase-accessible chromatin with sequencing to profile cortical and subcortical regions in post-mortem brain-tissue samples from Latin, white (excluding Latin) and African American (excluding Latin) individuals. Using discrete and continuous dissections of molecular programs, we identify cell-type-specific clusters associated with Alzheimer's disease in a region-specific manner across all three population groups, including microglial (GPNMB+ and CD74+ subgroups), astrocytic (SERPINH1+, CD44+ and WIF1+ subgroups) and neuronal (SST+ GABAergic and superficial-layer glutamatergic) signatures. We also report continuous gene-expression factors in astrocytes and oligodendrocytes that are not captured by discrete cluster assignments, but which show strong associations with disease phenotypes; these factors are enriched for genes associated with annotated functions such as lipid processing and neurotransmitter reuptake. Finally, we find that molecular programs reveal six distinct&#xa0;subgroups of&#xa0;individuals with cognitive impairment that span all three populations, are not captured by neuropathology, and are instead distinguished by molecular&#xa0;signatures that are not universally present but are nonetheless associated with ante-mortem impairment. Overall, our study identifies key cell types and gene programs implicated in Alzheimer's disease that are shared across population groups, and underscores how representative sampling can capture both shared signatures and disease heterogeneity, thereby enabling better prioritization of key cell types for further investigation.

Female↗

Chromosome-level genome assembly of Triplophysa scleroptera.

Triplophysa scleroptera is an endemic fish species in Qinghai Lake and the upper reaches of the Yellow River. However, studies on conservation and evolutionary genetics were seriously impeded by the absence of a reference genome. Here, by using PacBio HiFi sequencing and Hi-C assembly technology, we assembled a chromosome-level genome of T. scleroptera, with a total length of 660.22&#x2009;Mb and 99.82% of the sequence anchored to 25 chromosomes. The contig N50 and scaffold N50 were 9.09&#x2009;Mb and 24.38&#x2009;Mb, respectively. The evaluation using BUSCO indicated the genome assembly to be 96.40% complete. About 33.41% of the genome consists of repeat elements. We predicted 26,168 protein-coding genes in the genome, and 99.02% of them were functionally annotated. This high-quality reference genome would serve as a valuable genomic resource for advancing evolutionary conservation genetics studies in this species.

Animals↗

Chromosome-level genome assembly with telomeric repeats at scaffold ends for Rhabdosargus sarba.

Rhabdosargus sarba, the goldlined seabream, is a euryhaline marine fish of great aquaculture potential. Genome sequencing and assembly of R. sarba was carried utilizing a multi-platform sequencing strategy that included long-read sequencing (PacBio HiFi), short-read sequencing (Illumina), and chromatin interaction mapping (Hi-C). The final genome assembly size after scaffolding was 764.59&#x2009;Mb in 31 scaffolds with an N50 length of 33.98&#x2009;Mb. Repeat profiling of primary assembly showed that 28.71% of the genome comprises of repeat elements. Gene prediction utilising the evidence from ab initio prediction and transcriptome data revealed 26,913 protein encoding genes and functional annotation and pathway analysis showed their participation in 332 pathways. This genome is an excellent resource for future research on genetic improvement and molecular breeding programmes for R. sarba.

Animals↗

Chromosome-level genome assembly of Ceroplastes pseudoceriferus Green, 1935 (Hemiptera: Coccidae).

Soft scales (Hemiptera: Coccidae) are significant polyphagous pests and majority of which are invasive species. The 364.14&#x2009;Mb chromosome-level genome of Ceroplastes pseudoceriferus was assembled in this work, with a contig N50 length of 6.16&#x2009;Mb and scafold N50 length of 21.24&#x2009;Mb. Approximately 99.89% of assembled sequences were anchored into 18 chromosomes with the assistance of Hi-C reads. Furthermore, approximately 53.98% of the genome was composed of repetitive elements. In total, 10,475 protein-coding genes were predicted, of which 9503 (90.72%) genes were functionally annotated. The BUSCO analysis demonstrated the completeness of the genome annotation is 92.54%. This genome represents first high-quality chromosome level assembly of Coccidae, thereby advancing our knowledge of Coccidae insects and developing effective management strategies that protect crops, forests, and natural ecosystems.

Animals↗

Chromosome-level genome assembly of the longhorn beetle Arhopalus rusticus (Coleoptera: Cerambycidae).

The longhorn beetle Arhopalus rusticus (Coleoptera: Cerambycidae) is a widely distributed wood-boring pest of conifers. Here, we assembled a chromosome-level genome of A. rusticus using Illumina, Oxford Nanopore, and Hi-C sequencing technologies. The assembled genome is 1180.40&#x2009;Mb, with a scaffold N50 of 125.01&#x2009;Mb, and BUSCO completeness of 93.6%. All contigs were assembled into ten pseudo-chromosomes. The genome contains 69.87% repeat sequences. We identify 18, 377 protein-coding genes in the genome, of which 11,368 were functionally annotated. This genome provides a valuable resource for understanding the ecology, genetics, and evolution of A. rusticus, as well as for controlling wood-boring pests.

Animals↗

Chromosomal level genome assembly of medicinal plant Chrysosplenium macrophyllum.

Chrysosplenium macrophyllum Oliv., a perennial herb native to China, is widely used in traditional medicine for its notable therapeutic properties. However, the absence of a reference genome has constrained its full potential for research and application. This study presents the first chromosome-level de novo genome assembly of C. macrophyllum, constructed by integrating long reads from Oxford Nanopore Technologies (ONT), short reads from BGI, and Hi-C data. The final assembly spans 2.55&#x2009;Gb, with a scaffold N50 of 93.38&#x2009;Mb, and 83.70% of the genome has been assigned to 22 chromosomes. The mapping rate of the BGI short reads to the genome is approximately 97.94%, and BUSCO analysis reveals that 97.94% of the predicted genes are complete. A total of 62,921 protein-coding genes were predicted, with functional annotations for 93.67% of them. This chromosome-level genome assembly represents an important resource for expanding our understanding of Chrysosplenium species and supports future genomic studies and applications.

Genome, Plant↗

Near-complete reference genome assembly of Hoya carnosa.

Hoya R. Br. is the largest genus in the tribe Marsdenieae (Apocynaceae), comprising 350-450 species. Hoya species are popular in horticulture for their distinctive floral traits and fragrances, primarily sourced from domestication and mutation breeding. However, the lack of molecular analysis for floral morphological traits has limited their cultivation and application. In this study, we assembled a near-complete reference genome for H. carnosa, the model species of the genus, using PacBio HiFi reads and Hi-C method. The genome size was approximately 465.7&#x2009;Mb with a contig N50 of 39.3&#x2009;Mb. 99.7% of the sequences were anchored to 11 pseudochromosomes, and the assembly achieved a BUSCO score of 98.5%. We predicted 24,309 protein-coding genes, of which 90.2% (21,927) were functionally annotated. This high-quality genome provides a valuable reference for the research of evolution, conservation and molecular breeding in Hoya.

Genome, Plant↗