Search PubMedSearch

SEARCH · Search PubMed

Results for “Protein aggregation databases”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

9 recordsLinked to original sources

Prediction and Evaluation of Protein Aggregation with Computational Methods.

Protein and peptide aggregation has recently become one of the most studied biomedical problems due to its central role in several neurodegenerative disorders and of biotechnological importance. Multiple in silico methods, databases, tools, and algorithms have been developed to predict aggregation of proteins and peptides to better understand fundamental mechanisms of various aggregation diseases. Here, we attempt to provide a brief overview of bioinformatic methods and tools to better understand molecular mechanisms of aggregation disorders. Furthermore, through a better understanding of protein aggregation mechanisms, it might be possible to design novel therapeutic agents to treat and hopefully prevent protein aggregation diseases.

Computational Biology

Proteomics identify disease-associated variants in patients with rare diseases undiagnosed after genome sequencing.

Despite the introduction of genome sequencing (GS) for rare disease diagnostics, a genetic cause is not identified in most patients. Here, we explored the potential of proteomics to improve the diagnostic yield in 424 patients with rare diseases from the 100,000 Genomes Project (100kGP) without a genetic diagnosis. Serum proteomic profiling was performed using the Olink Explore 1536 assay (N&#xa0;=&#xa0;1463 proteins). For 13 patients without genetic diagnoses, detection of lower serum protein "outliers" (z-score&#xa0;<&#xa0;-2) led to confirmed genetic diagnoses by resolving variants of uncertain significance or prioritizing genes for targeted GS reanalysis. For 23 additional patients without genetic diagnoses (64% of findings), we identified candidate gene-disease links and variants through convergent evidence from lower protein outliers and variants ranked through the variant prioritization tool Exomiser. For example, we identified a candidate heterozygous missense variant [Genome Aggregation Database (gnomAD) minor allele frequency&#xa0;=&#xa0;0.006%] in tyrosine kinase with immunoglobulin-like and epidermal growth factor homology domains 1 (TIE1) that was only present in a patient with lower TIE1 serum abundance (z-score&#xa0;=&#xa0;-5.12) and their father, both of whom were affected by the same monogenic cardiac disorder, but in no other individuals from the 100kGP. Missense (52.5%) and splice region (27.5%) variants accounted for most diagnostic or candidate variants prioritized. This proof-of-principle study demonstrated that serum proteomics can support rare disease diagnosis and identify disease-causing genes in patients undiagnosed after GS, although successful implementation will likely depend on tissue specificity of protein expression, detectability in blood, proteomic platform coverage, and sensitivity.

Humans

Genome-Wide and Rare Variant Association Studies of Amblyopia in Admixed American and African Ancestry Groups.

OBJECTIVE: To identify genetic variants associated with amblyopia in African (AFR) and Admixed American (AMR) ancestry groups, expanding on previous studies conducted in European ancestry. DESIGN: Retrospective ancestry-stratified genome-wide association study (GWAS) and gene-level rare variant association study (RVAS). PARTICIPANTS: Participants in the All of Us Research Program from AFR and AMR ancestry groups who had whole-genome sequencing available. Cases and controls were distinguished based on the presence of International Classification of Diseases 9/10/SNOMED diagnosis codes for amblyopia in electronic health records. This yielded ancestry-stratified subsets of 269 cases and 71 585 controls of AMR ancestry and 366 cases and 79 460 controls of AFR ancestry. METHODS: Stratified logistic regression models were adjusted for age, biological sex, and the top 10 principal components of genomic ancestry. GWAS was limited to common variants (minor allele frequency &#x2265;1%), and RVAS was limited to rare variants with coding sequence-altering effects (minor allele frequency >1%, exonic only, excluding synonymous variants) aggregated at the gene level using the SKAT algorithm. Downstream analyses of the significant variants were performed using KEGG and GO pathway analysis and STRING database queries for protein-protein interactions and gene-gene interactions. MAIN OUTCOME MEASURES: Single-nucleotide polymorphisms were determined to have genome-wide significance if P < 5e-8 in the GWAS, and genes were determined to have significant association with amblyopia in the RVAS if P < 8.0 &#xd7; 10-4. RESULTS: In the AMR GWAS, 245 unique single-nucleotide polymorphisms mapping to 97 distinct loci were identified, notably within neurodevelopmental and axonal guidance genes, including ROBO1, SEMA4B, PTPRD, NRXN1, and CAMK2D. The AFR GWAS identified 11 significant variants corresponding to 6 loci mapping primarily to long noncoding RNAs and pseudogenes. The AMR RVAS identified 15 genes, including axonal transport genes (KIF1B and KIF7) and growth factor signaling genes (EGF, ERBIN, and AKAP17A). The AFR RVAS identified a single gene, DLG2, which encodes the postsynaptic protein PSD-93, which promotes the closure of the sensitive period of neuroplasticity for vision in early childhood. CONCLUSIONS: Genetic risk architectures for amblyopia differ across ancestries but fundamentally converge on neurodevelopmental signaling, cortical synapse assembly, and sensitive period plasticity rather than ocular structural dynamics. FINANCIAL DISCLOSURE(S): The authors have no proprietary or commercial interest in any materials discussed in this article.

Amblyopia

Knowledge-enhanced protein subcellular localization prediction from 3D fluorescence microscope images.

MOTIVATION: Pinpointing the subcellular location of proteins is essential for studying protein function and related diseases. Advances in spatial proteomics have shown that automatic recognition of protein subcellular localization from images could highly facilitate protein translocation analysis and biomarker discovery, but existing machine-learning works have been mostly limited to processing 2D images. By contrast, 3D images have higher spatial resolution&#xa0;and allow researchers to observe cellular structures in their natural context, but currently, there are only a few studies of 3D image processing for protein distribution analysis due to the lack of data and complexity of modeling. RESULTS: We developed a knowledge-enhanced protein subcellular localization model, KE3DLoc, which could recognize distribution patterns in 3D fluorescence microscope images using deep learning methods. The model designs an image feature extraction module that incorporates information from 3D and 2D projected cells and implements asymmetric loss and confidence weights to address data imbalance and weak cell annotation issues. Besides, considering that the biological knowledge in the Gene Ontology (GO) database can provide valuable support for protein location understanding, the KE3DLoc model incorporates a novel knowledge enhancement module that optimizes the protein representation by related knowledge graphs derived from the GO. Since the image module and the knowledge module calculate features from different levels, KE3DLoc designs protein ID aggregation to enhance the consistency of protein features across different cells. Experimental results on three public datasets have demonstrated that the KE3DLoc significantly outperforms existing methods and provides valuable insights for spatial proteomics research. AVAILABILITY AND IMPLEMENTATION: All datasets and codes used in this study are available at GitHub: https://github.com/PRBioimages/KE3DLoc.

Microscopy, Fluorescence

The molecular mechanism of cuproptosis and research progress in pancreatic diseases.

PURPOSE: Cuproptosis has been proven to be a novel mode of cell death, distinct from other types of cell death such as necrosis, ferroptosis, pyroptosis, and apoptosis. This study aims to systematically review the molecular mechanisms of cuproptosis in recent years and its research progress in pancreatic diseases. METHODS: By searching PubMed and Web of Science databases, 113&#x2009;key literatures were included for thematic analysis, covering the molecular mechanism of cuproptosis and its role in the occurrence and development of pancreatic cancer, acute and chronic pancreatitis, diabetes, pancreatic cyst, pancreatic injury and pancreatic neuroendocrine tumor. RESULTS: Cuproptosis refers to the accumulation of copper ions in cells, which leads to instability of ferritin and aggregation of acylated proteins, resulting in oxidative stress-related cell death. Recent studies have shown that cuproptosis plays an important role in the occurrence and development of various pancreatic diseases, such as pancreatic cancer, acute and chronic pancreatitis, diabetes, pancreatic cysts, pancreatic injuries and pancreatic neuroendocrine tumor. The inducers of cuproptosis, such as disulfiram, chloroquinolones, and perilla phenols, alleviate pancreatic cancer by promoting cell cuproptosis. Copper chelators such as tetraethylenepentamine and tetrathiomolybdate promote the recovery of pancreatic injury by inhibiting cell cuproptosis. CONCLUSIONS: Cuproptosis plays a crucial role in the pathogenesis of pancreatic diseases. Further research on the cuproptosis pathway may become a potential target for the treatment of pancreatic diseases.

Animals

Microbiome Datahub: an open-access platform integrating environmental metadata, taxonomy, and functional annotation for comprehensive metagenome-assembled genome datasets.

BACKGROUND: Metagenome-assembled genomes (MAGs) provide crucial insights into the genomic diversity of uncultured microbes. However, MAG datasets deposited in public repositories such as INSDC are often difficult to reuse due to heterogeneous quality, inconsistent taxonomic and functional annotations, and insufficiently curated environmental metadata. While secondary MAG databases such as MGnify, IMG/M, and SPIRE provide standardized resources, they reconstruct MAGs de novo from public metagenomic reads and therefore do not represent the original MAGs reported in publications. RESULTS: To address this gap, we developed Microbiome Datahub, an open-access platform that systematically aggregates and re-annotates original MAGs from INSDC. We collected 214,427 MAGs, predicted genes by DFAST, performed quality assessment with CheckM, standardized taxonomic assignments with GTDB-Tk, inferred 27 phenotypic traits using Bac2Feature, assigned proteins to MBGD ortholog clusters and KEGG Orthology IDs using PZLAST, and annotated environmental metadata with the Metagenome and Microbes Environmental Ontology. Across these MAGs, the average completeness was 80.5% and contamination 1.8%; notably, the most frequent values were&#x2009;>95% completeness and&#x2009;<1% contamination, indicating that the majority of MAGs are of high quality. Comparative analyses showed that Microbiome Datahub provides phylogenetically and environmentally diverse MAGs: while the majority originated from vertebrate gut environments, a substantial number were also recovered from other habitats such as groundwater, including nearly 10,000 MAGs from the Patescibacteria. Inference of 27 phenotypic traits, including optimum growth temperature, further revealed ecological differentiation across phyla. Protein clustering revealed 56 million identity 40% clusters, with the majority unique compared with MGnify and GlobDB, and&#x2009;~19% of proteins unassigned to MBGD ortholog clusters, underscoring their novelty. CONCLUSIONS: Microbiome Datahub integrates MAG genome sequences, gene and protein predictions, quality metrics, environmental and taxonomic annotations, ortholog cluster assignments, and phenotype predictions, all accessible via a web interface, API, and bulk downloads. By combining original MAGs with curated metadata and functional annotations, Microbiome Datahub constitutes a comprehensive and reusable resource that will accelerate microbiome and microbial genomics research. Video Abstract.

Metagenome

Unique signatures of highly constrained genes across publicly available genomic databases.

PURPOSE: Publicly available genomic databases are critical in understanding human genetic variation. They also provide unique insights into patterns of genetic constraints and their relationship with human disease. METHODS: We utilized one of the largest publicly available databases, Genome Aggregate Database, to determine genes that are highly constrained for only loss-of-function, only missense, and both loss-of-function/missense variants. We identified their unique signatures and explored their causal relationship with human diseases. Those genes were also evaluated for chromosomal location, tissue-level expression, Gene Ontology analysis, and gene family categorization using multiple publicly available databases. RESULTS: We identified unique patterns of inheritance, protein size, and enrichment in distinct molecular pathways for those constrained genes associated with human disease. In addition, we identified genes that are currently not known to cause human disease, which may be excellent gene discovery candidates. CONCLUSION: We elucidate biological pathways of highly constrained genes that expand our understanding of critical cellular proteins. The findings can also advance research in rare diseases.

Humans

pLAST-a tool for rapid comparison and classification of bacterial plasmid sequences.

MOTIVATION: The increasing number of fully sequenced bacterial plasmids being annotated and catalogued has prompted the development of computational tools for comparing and classifying them. Existing approaches typically compare full-length DNA sequences (e.g. Mash, BLASTn, and ANI-based methods) or translated open reading frames (ORFs) (e.g. DIAMOND), with plasmid-level scores obtained by aggregating ORF-to-ORF similarities; however, they are either restricted to closely related plasmids or become computationally demanding in large-scale analyses. RESULTS: We describe pLAST (plasmid Language Analysis and Search Tool), a plasmid-search tool built using word2vec representations of protein-family content informed by local genomic context. Benchmarks indicate that pLAST outperforms nucleotide-based methods and performs comparably to DIAMOND in identifying functionally similar plasmids and compared with the widely used Mash, it achieves 26% and 24% improvements in detecting shared mating-pair formation system type and relaxase type, respectively. This performance scales to database searches across hundreds of thousands of sequences, as demonstrated using the precomputed PlasmidScope collection of &#x223c;750&#xa0;000 plasmids. Beyond global similarity, pLAST also returns per-ORF plasmid-plasmid alignments, enabling detection of shared functional modules. AVAILABILITY AND IMPLEMENTATION: pLAST is freely accessible as a web server at&#x202f;https://plast.lbs.cent.uw.edu.pl/ or https://plast.lbs.biol.uw.edu.pl/ and available as a Python module along with a precomputed database at&#x202f;https://github.com/labstructbioinf/pLAST for customized analysis.

Plasmids

RNA/DNA Binding Protein TDP43 Regulates DNA Mismatch Repair Genes with Implications for Genome Stability.

TDP43 is an RNA/DNA binding protein increasingly recognized for its role in neurodegenerative conditions, including amyotrophic lateral sclerosis and frontotemporal dementia (FTD). As characterized by its aberrant nuclear export and cytoplasmic aggregation, TDP43 proteinopathy is a hallmark feature in over 95% of ALS/FTD cases, leading to the formation of detrimental cytosolic aggregates and a reduction in nuclear functionality within neurons. Building on our prior work linking TDP43 proteinopathy to the accumulation of DNA double-strand breaks (DSBs) in neurons, the present investigation uncovers a novel regulatory relationship between TDP43 and DNA mismatch repair (MMR) gene expressions. Here, we show that TDP43 depletion or overexpression directly affects the expression of key MMR genes. Alterations include MLH1, MSH2, MSH3, MSH6, and PMS2 levels across various primary cell lines, independent of their proliferative status. Our results specifically establish that TDP43 selectively influences the expression of MLH1 and MSH6 by influencing their alternative transcript splicing patterns and stability. We furthermore find aberrant MMR gene expression is linked to TDP43 proteinopathy in two distinct ALS mouse models and post-mortem brain and spinal cord tissues of ALS patients. Notably, MMR depletion resulted in the partial rescue of TDP43 proteinopathy-induced DNA damage and signaling. Moreover, bioinformatics analysis of the TCGA cancer database reveals significant associations between TDP43 expression, MMR gene expression, and mutational burden across multiple cancers. Collectively, our findings implicate TDP43 as a critical regulator of the MMR pathway and unveil its broad impact on the etiology of both neurodegenerative and neoplastic pathologies.

Amyotrophic lateral sclerosis