Search PubMedSearch

SEARCH · Search PubMed

Results for “protein structure prediction”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Uncovering viral protein acquisition events and human-specific folds with pairwise comparisons of predicted protein structures.

Pairwise sequence comparisons are at the center of molecular evolutionary analyses. However, viral pairwise comparisons are challenging because extreme mutation rates and evolutionary pressure cause genomes to diverge rapidly, limiting detectable sequence similarity to fewer than 3% of virus pairs. To overcome these limitations, we compared viruses based on structural similarity, using predicted protein structures from ColabFold and Foldseek to define protein fold clusters. We represented each virus genome by its protein structural content. Pairwise similarities between viruses were then quantified using the Jaccard index based on the presence or absence of protein fold clusters. Using a recently established viral protein fold database, we compared all pairs of eukaryotic viruses in RefSeq. This approach increased the proportion of comparable viral genome pairs from 2.4% to 16.5%. Using this protein-fold representation of viruses, we were able to accurately predict viral families with an average sensitivity of 85.9%. Investigation of viral families showing limited sensitivity with this approach uncovered a laterally transferred structural cluster (Rep/NS1) broadly shared across diverse viral families and found in the avian lineage of adenoviruses. Sequence homology suggests that this Rep was acquired from Parvoviridae, but the protein is mutant in the ATPase active site, indicating possible exaptation toward a purely DNA-binding function. In Gammapapillomaviruses, several E4 clusters were associated with human tropism. In summary, by representing viruses with structural protein clusters, we can classify highly divergent viruses, trace lateral gene transfer, and uncover features associated with viral host range.

Humans

CASTER-DTA: Equivariant Graph Neural Networks for Predicting Drug-Target Affinity.

Accurately determining the binding affinity of a ligand with a protein is important for drug design, development, and screening. With the advent of accessible protein structure prediction methods such as AlphaFold, predicted protein 3D structures are readily available; however, methods for predicting binding affinity currently do not take full advantage of 3D protein information. Here, we present CASTER-DTA (Cross-Attention with Structural Target Equivariant Representations for Drug-Target Affinity), which uses an equivariant graph neural network to learn more robust protein representations alongside a standard graph neural network to learn molecular representations to predict drug-target affinity. We augment these representations by incorporating an attention-based mechanism between protein residues and drug atoms to improve interpretability. We show that CASTER-DTA represents a state-of-the-art improvement on multiple benchmarks for predicting drug-target affinity and that it generates novel insights for several related tasks. We then apply CASTER-DTA to create a large resource of the binding affinities of every FDA-approved drug against every protein in the human proteome and make these predictions freely available for download. We also make available a web server for researchers to apply a pretrained CASTER-DTA model for predicting binding affinities between arbitrary proteins and drugs.

deep learning

ProtPen Combines Sequence- and Structure-based Approaches to Facilitate Protein Function Predictions on a Proteome-wide Scale.

Proteins of unknown function represent a significant gap in our understanding of biological processes, encompassing large portions of the proteomes of many organisms, especially prokaryotes. Addressing this gap is critical to understanding the biology and pathogenicity of such organisms. We introduce ProtPen, an open-source pipeline that facilitates protein function prediction by combining eggNOG-mapper for sequence-based annotation with Foldseek for rapid structural similarity searches using AlphaFold-predicted protein structures. Annotation results from both tools are merged and enriched with UniProt metadata to produce a comprehensive output suitable for downstream analysis. The pipeline requires only a FASTA input file with UniProt identifiers, and is designed to analyze data sets on the scale of whole proteomes. Benchmarking on a curated data set of well-characterized Pseudomonas aeruginosa proteins demonstrated an annotation accuracy of >90%, and highlighted the complementarity of sequence- and structure-based methods. Further evaluation of ProtPen included its application to biologically relevant data sets, comprising proteins of unknown function that exhibited significant differential abundances in a proteomics data set of P. aeruginosa, and uncharacterized glycoproteins from Haloferax volcanii. ProtPen is readily extensible to incorporate additional protein function prediction tools. In summary, this pipeline facilitates the systemwide annotation of proteins of unknown function from proteomic data sets and whole proteomes.

Pseudomonas aeruginosa

Recurrent Evolutionary Innovations in Rodent and Primate Schlafen Genes.

SCHLAFEN proteins are a large family of RNase-related enzymes carrying essential immune and developmental functions. Despite these important roles, Schlafen genes display varying degrees of evolutionary conservation in mammals. While this appears to influence their molecular activities, a detailed understanding of these evolutionary innovations is still lacking. Here, we used in-depth phylogenomic approaches to characterize the evolutionary trajectories and selective forces shaping mammalian Schlafen genes. We traced lineage-specific Schlafen amplifications and found that recent duplicates evolved under distinct selective forces, supporting repeated subfunctionalization cycles. Codon-level natural selection analyses in primates and rodents identified recurrent positive selection over Schlafen protein domains engaged in viral interactions. Combining known crystal structures and predicted protein structures, we discovered a novel class of rapidly evolving residues enriched at the contact interface of SCHLAFEN protein dimers. Our results suggest that inter-SCHLAFEN compatibilities are under strong selective pressures and are likely to impact their molecular functions. We posit that cycles of genetic conflicts with pathogens and between paralogs drove Schlafens' recurrent evolutionary innovations in mammals.

Animals

Structural genomics sheds light on protein functions and remote homologs across the insect tree of life.

Protein structure bridges the sequence-function relationship, enabling deep exploration of biological processes across diverse organisms. Insects, the most diverse animal lineage, accounting for over 50% of all described animal species, provide an exceptional system for exploring sequence-structure-function relationships. Here, we reconstructed a comprehensive and well-resolved phylogeny of 4854 insects, spanning all orders. Leveraging this framework, we created an atlas of 13.29 million predicted protein structures from 824 representative species, including 11.63 million newly predicted structures. Structural clustering revealed that proteins with divergent sequences but similar structures could be effectively grouped together. Structural similarity searches against proteins with well-characterized functions yielded annotations for 7.61 million insect proteins, including up to 14% of previously unannotated proteins. We further identified 750 million remote homologs between insect proteins, many of which trace back to ancient branches of the insect phylogeny. Remarkably, despite extensive sequence divergence, cGAS-like receptors (cGLRs) were structurally conserved across all 824 insects. Experimental assays demonstrated that these structurally identified cGLRs play a crucial role in antiviral defense in the yellow fever mosquito. Our findings highlight the significance of structural genomics for understanding protein function and evolution across the tree of life.

Animals

Improving RNA Secondary Structure Prediction Through Expanded Training Data.

In recent years, deep learning has revolutionized protein structure prediction, achieving remarkable speed and accuracy. RNA structure prediction, however, has lagged behind. Although several methods have shown some success in predicting RNA secondary and tertiary structures, none have reached the accuracy observed with contemporary protein models. The lack of success of these RNA structure prediction models has been proposed to be due to limited high-quality structural information that can be used as training data. To probe this proposed limitation, we developed a large and diverse dataset comprising paired RNA sequences and their corresponding secondary structures. We assess the utility of this enhanced dataset by retraining on a deep learning model, SincFold. We find that SincFold exhibited improved generalization to some previously unseen RNA families, enhancing its capability to predict accurate de novo RNA secondary structures. The RNASSTR dataset provides a substantial advance for RNA structure modeling, laying a strong foundation for the development of future RNA secondary structure prediction algorithms.

Journal Article

Expanding kinetoplastid genome annotation through protein structure comparison.

Kinetoplastids belong to the Discoba supergroup, an early divergent eukaryotic clade. Although the amount of genomic information on these parasites has grown substantially, assigning gene functions through traditional sequence-based homology methods remains challenging. Recently, significant advancements have been made in in-silico protein structure prediction and algorithms for rapid and precise large-scale protein structure comparisons. In this work, we developed a protein structure-based homology search pipeline (ASC, Annotation by Structural Comparisons) and applied it to transfer biological information to all kinetoplastid proteins available in TriTrypDB, the reference database for this lineage. Our pipeline enabled the assignment of structural similarity to a substantial portion of kinetoplastid proteins, improving current knowledge through annotation transfer. Additionally, we identified structural homologs for representatives of 6,700 uncharacterized proteins across 33 kinetoplastid species, proteins that could not be annotated using existing sequence-based tools and databases. As a result, this approach allowed us to infer potential biological information for a considerable number of kinetoplastid proteins. Among these, we identified structural homologs to ubiquitous eukaryotic proteins that are challenging to detect in kinetoplastid genomes through standard genome annotation pipelines. The results (KASC, Kinetoplastid Annotation by Structural Comparison) are openly accessible to the community at kasc.fcien.edu.uy through a user-friendly, gene-by-gene interface that enables visual inspection of the data.

Kinetoplastida

[Genetic analysis of a male with Multiple morphological abnormalities of sperm flagella combined with sperm head abnormalities due to compound heterozygous variants of DNAH1 gene and a literature review].

OBJECTIVE: To explore the clinical phenotype and genetic etiology of a male with Multiple morphological abnormalities of sperm flagella (MMAF) combined with sperm head abnormalities due to compound heterozygous variants of DNAH1 gene, with an aim to provide guidance for assisted reproductive technology in his family. METHODS: A man with MMAF combined with sperm head abnormalities who visited Women and Children's Hospital of Ningbo University in October 2024 was selected as study subject. Clinical data of the patient's family were retrospectively collected. Peripheral blood samples were collected from the patient and his spouse, and G-banding karyotyping and whole exome sequencing (WES) were carried out. Candidate variants were validated by Sanger sequencing. Conservation of the DNAH1 protein was queried on the UCSC website. The difference between wild type and variant DNAH1 proteins were analyzed using AlphaFold v3.0.1 and PyMOL v2.5.6. The pathogenicity of variant was rated based on the guidelines from American College of Medical Genetics and Genomics (ACMG). Previous literature was searched using keywords "DNAH1 gene" and "multiple morphological abnormalities of the sperm flagella" on CNKI, Wanfang Data Knowledge Service Platform, and PubMed database to identify cases of MMAF attributed to biallelic DNAH1 gene variants. The retrieval period was set from the establishment of the databases to December 31, 2025. The genotypes and clinical phenotypes of patients with biallelic DNAH1 mutations were analyzed. This study was approved by the Medical Ethics Committee of the hospital (Ethics No.: EC2023-094). RESULTS: The 30-year-old patient and his 30-year-old wife had infertility for 2 years. Semen analysis revealed no motile sperm and a 99.0% abnormal morphology rate. Typical MMAF was observed with phase-contrast microscopy. Sperm morphology analysis revealed abnormalities of the head, neck, and tail with an approximate ratio of 9:5:1. The patient's karyotype was 46,XY, and his wife's karyotype was 45,X[4]/47,XXX[1]/46,XX[84]. WES and Sanger sequencing revealed that the patient harbored compound heterozygous variants of the DNAH1 gene, namely c.1435_1444+3del and c.12204_12206del (p.Asn4069del), but their origin remained unidentified. UCSC genome browser query results showed that the amino acid residue at position 4 069 of the DNAH1 protein is highly conserved across various species. Protein structure prediction reveals that, in the wild-type DNAH1 protein, the Asparagine at position 4 069 (Asn4069) can form hydrogen bonds with the Leucine on the main chain at position 4 086 (Leu4086) and the Serine on the side chain at position 4 087 (Ser4087). The c.12204_12206del variant, resulting in deletion of Asn4069, disrupts these hydrogen bonds and does not generate any compensatory interactions. Based on the ACMG guidelines, the c.1435_1444+3del variant was predicted to be likely pathogenic (PM2_Supporting+PVS1), and the c.12204_12206del(p.Asn4069del) variant was rated as likely pathogenic (PM2_Supporting+PM4+PM3+PP4). The couple had elected for in vitro fertilization using donor sperm. During this cycle, 12 oocytes were retrieved, 10 oocytes were successfully fertilized, 1 embryo and 6 blastocysts were obtained. Following the first transfer of a frozen-thawed blastocyst, implantation of an empty gestational sac occurred, which led to a miscarriage. After the second transfer of a high-quality blastocyst, the embryo split into twins following implantation, and the spouse had selected fetal reduction. The gestational age was 33+3 weeks on June 1, 2026. Literature review identified three studies reporting biallelic mutations of the DNAH1 gene in association with MMAF combined with sperm head abnormalities. Together with the patient from this study, a total of 20 patients were included in the analysis. The rate of sperm flagellar abnormalities in these patients was above 80.0%, while the rate of sperm head abnormalities has ranged from 12.0% to 100.0%. In four patients, the genetic basis was unknown. In the remaining 16 patients, 35 mutations were detected, with c.8626-1G>A being the most common (22.9%, 8/35). CONCLUSION: This patient showed MMAF with frequent sperm head defects. Compound heterozygous variants of the DNAH1 gene probably underlay these abnormalities, which in turn has led to his primary infertility. This study revealed the phenotypic variability of MMAF and broadened the mutational spectrum of the DNAH1 gene.

Humans

The emergence of putative epistatic mutations and iSNVs in SARS-CoV-2 XBB.1.16 variants linked with alteration in immunogenic determinants.

The SARS-CoV-2 XBB variants have been proposed to evolve towards immune evasion against vaccination or natural infection, which may contribute to higher transmissibility. The XBB.1.16 independently emerged due to accumulation of two important substitutions, E180V and T478R in the spike protein. Its pseudoviral infectivity and evasion of humoral immunity were similar to XBB.1 and XBB.1.5. In March 2023, XBB.1.16 had outcompeted other dominant XBB variants in India, which indicate a potential growth advantage. Here, intra-host single nucleotide variations (iSNV) and mutations were screened in SARS-CoV-2 genomes in closely related individuals at two time points: at symptoms onset, and during recovery. The prominence of putative epistatic iSNVs (E180V, G184V, G252V, D253G, and P521S/T) in XBB.1.16 variants were detected during the recovery phase. E180V exhibits mutational constellations with the G252V and P521T in a subset of samples, and this pattern was also detected in contemporary SARS-CoV-2 genomes. Higher order protein structural predictions suggested that the putative epistatic interactions among E180V, G184V, and G252V, D253G may be associated with S protein folding and structural stability. This study involving genomics and computational analyses highlights the potential role of these putative epistatic interactions in immune evasion, which may have contributed to dominance of XBB variants.

Humans

[Analysis of a Chinese pedigree affected with Townes-Brocks syndrome due to a novel variant of SALL1 gene and a literature review].

OBJECTIVE: To analyze a novel exonic variant of the SALL1 gene and its impact on the binding site of SALL protein. METHODS: Clinical data of three children diagnosed with Townes-Brocks syndrome and their family members who had presented at the First Affiliated Hospital of Shandong First Medical University in April 2022 were retrospectively collected. The pathogenic variant was identified through whole-genome sequencing (WGS) and validated by Sanger sequencing. Protein structural prediction was performed using AlphaFold and PyMOL software to construct three-dimensional models of the wild-type and mutant proteins. Additionally, previously reported cases were systematically reviewed. This study was approved by the Medical Ethics Committee of the hospital (Ethics No.: 2023-386). RESULTS: The proband was one of triplet sisters born at 34+4 gestational weeks. All three cases had presented with anal atresia and rectovaginal fistula, and case 3 also had toe malformation of left foot. WGS revealed a novel heterozygous c.757C>T (p.Gln253*) variant in the SALL1 gene, which was predicted to be pathogenic. Sanger sequencing confirmed co-segregation of the variant with the disease within the family. Protein structural modeling demonstrated that the variant has introduced a premature stop codon at position 253, resulting in a truncated protein. CONCLUSION: Above finding has enriched the mutation spectrum of the SALL1 gene in association with Townes-Brocks syndrome, which also represented a rare case of anal atresia in triplets, and provided a basis for molecular diagnosis, genetic counseling, and further research.

Humans

MEG3 Promoter Methylation and F11 Receptor (F11R) Overexpression Define a High-Risk Subtype of Diabetic Pancreatic Cancer.

Long-standing diabetes mellitus (long-DM) (≧3 years) is associated with worse clinical outcomes in patients with pancreatic ductal adenocarcinoma (PDAC). Emerging evidence suggests that epigenetic alterations may contribute to this association; however, the underlying mechanisms remain largely unclear. This study aimed to elucidate the role of the tumor-suppressive long noncoding RNA maternally expressed gene 3 (MEG3) and related molecules in the development of PDAC with long-DM. A total of 117 patients who underwent surgical resection for PDAC at Hirosaki University Hospital were retrospectively analyzed. Histopathological assessment followed World Health Organization criteria and the Union for International Cancer Control tumor-node-metastasis classification. Promoter methylation of MEG3 was assessed via methylation-specific PCR using formalin-fixed paraffin-embedded tissue. MEG3 expression levels were assessed by real-time quantitative PCR. Additionally, proteomic profiling was performed using liquid chromatography-tandem mass spectrometry on formalin-fixed paraffin-embedded tissue samples. Among the 117 cases with PDAC, patients with long-DM exhibited significantly poorer tumor differentiation and reduced cancer-specific survival. MEG3 promoter methylation was more prevalent in patients with long-DM. MEG3 methylation was correlated with reduced MEG3 expression, increased venous invasion, higher recurrence rates, and worse prognosis. Proteomic analysis and protein structure prediction tool revealed F11 receptor (F11R) as a potential downstream effector of MEG3. F11R protein expression levels were evaluated using semiquantitative immunohistochemistry. Higher F11R expression was observed in patients with long-DM, correlating with poor histologic differentiation and unfavorable outcomes. Patients with PDAC showing simultaneous MEG3 methylation and F11R high expression were more likely to have long-DM, with additive effects of these changes and tumor recurrence. Our results demonstrated that MEG3 and its potential downstream regulator, F11R, could be involved in PDAC progression, particularly in patients with long-DM. The findings underscore the clinical significance of epigenetic regulation in DM-related PDAC, suggesting novel targets, such as MEG3 and F11R, for potential therapeutic intervention.

Humans

Deep learning-based assessment of missense variants in the COG4 gene presented with bilateral congenital cataract.

OBJECTIVE: We compared the protein structure and pathogenicity of clinically relevant variants of the COG4 gene with AlphaFold2 (AF2), Alpha Missense (AM), and ThermoMPNN for the first time. METHODS AND ANALYSIS: The sequences of clinically relevant Cog4 missense variants (one novel identified p.Y714F and three pre-existing p.G512R, p.R729W and p.L769R from Uniprot Q9H9E3) were imported into AF2 for protein structural prediction, and the pathogenicity was estimated using AM and ThermoMPNN. Different pathogenicity metrics were aggregated with principal component analysis (PCA) and further analysed at three levels (amino acid position, substitution and post-translation) based on all possible Cog4 missense variants (n=14 915). RESULTS: Localised protein structural impact including change of conformation and amino acid polarity, breakage of hydrogen bond and salt-bridge, and formation of alpha-helix were identified among clinically relevant Cog4 variants. The global structural comparison with multidimensional scaling demonstrated variants with similar protein structures (AF2) tended to exhibit similar clinical and biological phenotypes. The Cog4 p.Y714F variant exhibited greater protein structural similarity to mutated Cog4 found in Saul‒Wilson syndrome (p.G512R) and shared similar clinical phenotype (congenital cataract and psychomotor retardation). PCA of included pathogenic metrics demonstrated p.Y714F occurred at a critical position in Cog4 amino acid sequence with disrupted post-translational phosphorylation. CONCLUSION: Deep learning algorithms, including AF2, AM and ThermoMPNN, can be useful for evaluating variant of uncertain significance (VUS) by structural and pathogenicity prediction. Despite classified as VUS (American College of Medical Genetics and Genomics criteria: PM1, PP4), the pathogenicity in this Cog4 variant cannot be ruled out and warrants further investigation.

Mutation, Missense

Bioinformatic Analysis of Bacillus pacificus B630: Molecular Understanding of Biofilm Production.

The aim of this study was to determine biofilm production and motility in Bacillus pacificus B630 and Bacillus cereus ATCC 14579, and to perform a comparative genome analysis using bioinformatic tools to understand the differences between the two strains. Biofilm production was performed in glass tubes stained with safranin; motility was determined on soft agar. Bioinformatic analysis was performed using genomic information from both strains, including the identification of orthologous genes, the similarity between genes of the eps1 and sipW-tasA-calY operons, and the SipW and TasA model prediction. B. pacificus B630 produces a greater amount of biofilm on glass than B. cereus ATCC 14579 (p < 0.01). Furthermore, B. pacificus B630 shows lower motility than B. cereus ATCC 14579 (p < 0.001). B. pacificus B630 contains 45 unshared genes, whereas B. cereus ATCC 14579 has 27 unshared genes. Differences in similarity were observed between the genes of the eps1 and sipW-tasA-calY operons. These differences between SipW and TasA may affect protein structural predictions. In SipW, the differences may affect the C-terminal region. In TasA, the number of B-sheets differed between the two proteins, and amino acid substitutions were found in regions of high protein aggregation. Genomic differences in genes associated with biofilm production may explain differences in biofilm production between the strains studied.

Biofilms

Fast, accurate construction of multiple sequence alignments from protein language embeddings.

Multiple sequence alignment (MSA) is a foundational task in computational biology, underpinning protein structure prediction, evolutionary analysis, and domain annotation. Traditional MSA algorithms rely on pairwise amino acid substitution matrices derived from conserved protein families. While effective for aligning closely related sequences, these scoring schemes struggle in the low-identity "twilight zone." Here, we present a new approach for constructing MSAs leveraging amino acid embeddings generated by protein language models (PLMs), which capture rich evolutionary and contextual information from massive and diverse sequence datasets. We introduce a windowed reciprocal-weighted embedding similarity metric that is surprisingly effective in identifying corresponding amino acids across sequences. Building on this metric, we develop ARIES (Alignment via RecIprocal Embedding Similarity), an algorithm that constructs a PLM-generated template embedding and aligns each sequence to this template via dynamic time warping in order to build a global MSA. Across diverse benchmark datasets, ARIES achieves higher accuracies than existing state-of-the-art approaches, especially in low-identity regimes where traditional methods degrade, while scaling almost linearly with the number of sequences to be aligned. Together, these results provide the first large-scale demonstration of the power of PLMs for accurate and scalable MSA construction across protein families of varying sizes and levels of similarity, highlighting the potential of PLMs to transform comparative sequence analysis.

Deep Learning

Functional unknomics of the SAR11 clade reveal hidden genetic potential underlying adaptation to bottom-up and top-down pressures.

UNLABELLED: A substantial fraction of the genes in bacteria lack detectable sequence similarity to genes with known functions. These functionally uncharacterized genes-collectively referred to as the "unknome"-represent a largely unexplored genetic repertoire harboring insights into marine bacterial ecology. In this study, we explored the function of the unknome of the SAR11 clade, the most abundant bacterial lineage in the ocean, with a particular focus on genes that provide insight into its ecology. Based on the Clusters of Orthologous Genes and Kyoto Encyclopedia of Genes and Genomes classifications, approximately 56% of SAR11 ortholog groups were classified as members of the unknome. Among the SAR11 unknome, we successfully inferred the functions of 57 ortholog groups that are conserved in the SAR11 clade by protein structure similarity searches and genomic context analyses. These ortholog groups include putative transporter components, supporting the current ecological understanding that the SAR11 clade is specialized in substrate uptake to adapt to oligotrophic marine environments. Furthermore, structural analysis indicated that the DUF2237-containing protein, enriched in marine environments, may interact with purine nucleotide-containing compounds. This may suggest the existence of unique nucleotide utilization mechanisms in marine bacteria. In addition, we identified candidate viral defense systems within the unknome, indicating that diverse defense systems are present in at least one-third of cultured SAR11 strains. The presence of these defense systems, even within streamlined SAR11 genomes, suggests that they confer significant ecological advantages. Our analyses provide insights into the genetic basis of bottom-up processes (adaptation to oligotrophic environments) and top-down processes (antiviral defense strategy) contributing to ecological success. IMPORTANCE: Many microbial genes have no experimentally established function, limiting our ability to explain how microorganisms adapt to their environments. We examined this uncharacterized gene space, or "unknome" in SAR11, the most abundant bacterial clade in the ocean, by integrating evolutionary conservation, genomic context, predicted protein structure, and environmental distribution. This approach enabled us to prioritize components of the SAR11 unknome, including a core unknome conserved across the clade and genes enriched in specific lineages, and to identify several candidates with possible ecological roles in nutrient acquisition and defense against viruses. Our results suggest that the SAR11 unknome contains important clues to the ecological success of SAR11 rather than merely reflecting incomplete annotation or gene-prediction artifacts. Our study highlights the potential value of unknome analysis for identifying ecologically relevant genes in environmental microorganisms.

Pelagibacterales

A recurrent CCDC82 frameshift variant associated with syndromic neurodevelopmental disorder in a consanguineous Pakistani family.

BACKGROUND: Intellectual disabilities (IDs) are part of neurodevelopmental disorders (NDDs) and are genetically heterogeneous conditions characterized by impairments in cognition, learning, and adaptive functioning. Despite advances in gene discovery, many individuals, particularly those from understudied populations, remain without a molecular diagnosis. Recent reports implicate CCDC82 (HGNC: 26282) as an autosomal recessive ID gene, although the phenotypic spectrum and biological context remain incompletely defined. METHODS: Exome sequencing (ES) was performed in a consanguineous Pakistani family (PKMR06A) with four affected individuals presenting with moderate to severe ID. Variant segregation was confirmed by Sanger sequencing. In silico analyses, including pathogenicity prediction, protein structural modeling, and domain intolerance assessment, were used to evaluate the functional consequences of the identified variant. Spatiotemporal gene expression patterns were examined using bulk and single-cell human brain transcriptomic datasets. RESULTS: Clinically, affected individuals of family PKMR06A presented with early childhood global developmental delay, speech delay, hypotonia, gait abnormalities, spasticity, and mild facial dysmorphism. Genetic screening revealed a recurrent rare homozygous frameshift variant in CCDC82 (NM_024725.4): c.373del; p.(Asp125Ilefs*6), segregating with disease in all available affected individuals of the family. The identified c.373del variant was absent from the gnomAD database and was classified as pathogenic (PVS1, PM2, and PP1) based on ACMG/AMP criteria. The c.373del variant is predicted to introduce a premature termination codon, p.(Asp125Ilefs*6), leading to deletion of essential coiled-coil domains from the encoded protein, supporting a loss-of-function mechanism. In silico, transcriptomic analyses demonstrated preferential CCDC82 expression during prenatal human brain development, providing developmental context for the neurodevelopmental phenotype associated with the identified truncating variant. CONCLUSIONS: This study expands the mutational landscape of CCDC82 and provides additional clinical and molecular evidence supporting its role in autosomal recessive NDD. The findings reinforce the importance of CCDC82 in human neurodevelopment and highlight the value of genomic investigation in underrepresented populations.

Autosomal recessive

KSHVbook: An Information-Sharing Database for Kaposi's Sarcoma-Associated Herpesvirus.

Kaposi's sarcoma-associated herpesvirus (KSHV) is a double-stranded DNA virus belonging to the &#x3b3;-herpesvirus subfamily. KSHV is the causative agent of Kaposi's sarcoma (KS), primary effusion lymphoma (PEL), multicentric Castleman's disease (MCD), and KSHV inflammatory cytokine syndrome (KICS). Since its discovery, research on KSHV has rapidly progressed, but existing information platforms relatively lack comprehensiveness and do not provide efficient analysis tools tailored for KSHV. To further promote the research on KSHV more effectively, we have developed KSHVbook (http://www.kshvbook.com), a specialized information-sharing database dedicated to KSHV. This platform offers extensive information on genes, coding sequences, proteins, and the gene regulatory region. Besides, the KSHVbook includes about 35&#x2009;010 transcription factor binding sites (TFBSs), 342&#x2009;010 pairs of KSHV miRNA-host target gene relationships, protein structures predicted by AlphaFold3, qPCR primers, and so on. We also develop analytical tools for viral genome regions, TFBSs, and KSHV miRNA target genes to discover previously unknown biological functions of KSHV. These analytical tools can effectively identify the potential regulatory relationships between host transcription factors and viral genes. Overall, this platform provides a centralized data resource for KSHV research by integrating multiple databases, offering accessible analysis tools, and simplifying data acquisition. The KSHVbook will continue to be updated, and more features can be found on the website.

Herpesvirus 8, Human

Recent discovery of new enzymes in plant natural product biosynthesis.

Plants are a vast reservoir of natural products with diverse structural scaffolds, making them an invaluable source for discovering novel enzymes that catalyze unique and evolutionarily specialized metabolic transformations in biosynthetic pathways. Rapid advances in genomics, metabolomics, protein structure prediction, and heterologous pathway reconstruction have enabled the identification of numerous cryptic biosynthetic enzymes responsible for key scaffold-forming and tailoring reactions in metabolism. Particularly notable are the discoveries of plant-derived enzymes that catalyze challenging chemical transformations, including oxidative carbon-carbon bond rearrangements, atypical cycloadditions, radical-mediated coupling reactions, and iterative scaffold remodeling. This review summarizes major advances in enzyme discovery in plant natural product biosynthesis in recent years, focusing on emerging catalytic mechanisms, strategies for elucidating pathways, and evolutionary relationships, and highlights their implications for synthetic biology, metabolic engineering, and the sustainable production of valuable natural products.

Biological Products