Search PubMedSearch

SEARCH · Search PubMed

Results for “conservation score”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Functional Prediction of Epitranscriptome.

N6-methyladenosine (m6A) is one of the most prevalent and well-studied RNA modifications, playing a pivotal role in many biological processes. With the recent advances in high-throughput sequencing technologies, tens of thousands of m6A sites have been reported. However, not all m6A sites are important or functionally significant, highlighting the need to distinguish biologically relevant m6As from non-functional or technically artefactual ones. Here, we describe ConsRM, which is a web-based resource that was designed to evaluate the importance of m6As from an evolutionary perspective. It introduced a novel scoring framework for quantifying the conservation degree of m6As in humans. Its web interface includes a database of 177998 distinct human m6A sites along with their calculated conservation score, and allows users to analyze their own data via the web server. ConsRM is freely accessible at: http://180.208.58.19/conservation/browser.html .

Humans

GUANinE v1.1 reveals complementarity of supervised and genomic language models.

There has been much debate about the benefits of supervised versus unsupervised learning on genomes. Determining which is better in what contexts requires developing comprehensive benchmarks spanning functional and evolutionary tasks. Importantly, such benchmarks need large sample sizes to enable well-powered ranking of models. Having developed and applied such a benchmark here (GUANinE v1.1), we conclusively demonstrate each paradigm offers key advantages and outperforms on certain tasks. In accordance with training, supervised sequence-to-function models exhibit strong performance when annotating functional states characterized by chromatin accessibility or histone marks, while self-supervised language models outperform on evolutionary conservation. Our hundreds of new evaluations in this v1.1 expansion provide evidence for a tradeoff between input context size and model parameter count for a fixed compute budget, which we depict with new metrics such as kiloparameters/base pair. We also construct two new large-scale variant interpretation tasks in v1.1: cadd-snv measuring deleteriousness, and clinvar-snv measuring clinical pathogenicity. We find that conservation scores, and by extension, genomic language models, predict deleteriousness well, but successfully translating deleteriousness predictions to pathogenicity remains challenging. GUANinE v1.1 newly evaluates dozens of pretrained genomic models, and we conclude that moderate-context hybrid or post-trained language models may define the next era of machine learning in genomics.

Genomics

Effect of peer interaction on the problem-solving behavior of mentally retarded youths.

Educable mentally retarded young adults were given a one-bit logic problem (Experiment 1) or a test of conservation (Experiment 2) and classified as either low or high performers. In a second session the low performers were paired with high performers, and the dyads were required to agree on the solution to the problems presented. A control group of low performers was simply tested individually in a second session. One month later all the low performers were retested on the same logic problems or on an alternate form of the conservation test. No significant differences were observed in Experiment 1. In Experiment 2, the conservation scores of subjects who participated in the peer interaction were higher than those of controls on the posttest; however, based on responses to the lie item, we concluded that the experimental subjects had simply learned to respond "same" and to parrot a verbal statement. No such mimicking was possible in the logic problem. Results suggest that peer interaction is not a particularly effective method for enhancing the problem-solving ability of mentally retarded individuals.

Adult

A reinforcement learning-enhanced fuzzy multi-objective equilibrium optimization framework for multiple sequence alignment.

Multiple sequence alignment (MSA) is a fundamental task in bioinformatics, underpinning comparative genomics, structural analysis, and evolutionary inference. However, MSA remains a challenging multi-objective optimization problem due to the need to simultaneously maximize alignment accuracy, preserve conserved regions, and control gap proliferation, particularly in large and heterogeneous sequence collections. In this work, we propose MOFSACEO-MSA, a novel hybrid optimization framework for multiple sequence alignment that integrates a fuzzy multi-objective evaluation scheme with the Equilibrium Optimizer (EO) and a Soft Actor-Critic (SAC)-based adaptive control mechanism. The proposed framework formulates MSA as a dynamic multi-objective optimization problem, in which alignment quality is assessed using complementary residue-level and column-level criteria, including Sum-of-Pairs score, column conservation, entropy, and gap statistics. Fuzzy membership functions are employed to harmonize competing objectives into a unified optimization landscape, while EO provides robust global exploration. To further enhance adaptability, SAC dynamically regulates key EO parameters during the search process, enabling an effective balance between exploration and exploitation across datasets of varying size and heterogeneity. Extensive experiments werew conducted on diverse biological sequence datasets, with a primary focus on RNA benchmarks, including structured families from Rfam, large-scale repositories from RNAcentral and GenBank, and organism-specific tRNA datasets from GtRNAdb. Comparative evaluations against classical alignment tools (ClustalW, MAFFT, MUSCLE, PRANK, KAlign, and T-Coffee), metaheuristic methods (SAGA, Sequoya and EAFSA), and a reinforcement learning-based approach (RLALIGN) demonstrate that MOFSACEO-MSA consistently achieves competitive or superior Sum-of-Pairs scores while significantly reducing gap proportions and maintaining compact alignment lengths. Notably, the proposed framework exhibits improved robustness on large and highly heterogeneous datasets, where existing methods often suffer from excessive gap insertion or unstable convergence. Overall, MOFSACEO-MSA provides a flexible and extensible optimization paradigm that effectively bridges evolutionary search and reinforcement learning for high-quality multiple sequence alignment, with demonstrated effectiveness on challenging RNA alignment tasks.

Sequence Alignment

Oral hygiene and gingival health in Danish dental students and faculty.

This survey attempted to determine the status of oral cleanliness and gingival health in 150 dental students and 101 faculty members in a dental school. Without advance notice, plaque deposits were scored, using the Plaque Index System, and gingival health was determined using the criteria of the Gingival Index System. The 1st-year students had the poorest hygiene and gingival health. An improvement (P less than 0.01) was noted in the 2nd-year students who were still not in clinical training but had completed a course in preventive dentistry including oral hygiene techniques. Further improvement (P less than 0.05) was found in students participating in the clinical courses (3rd and 4th years). However, some deterioration of both hygiene and gingival status occurred in the senior 5th year. Among the faculty, the best oral hygiene and gingival state were found in members of departments in which clinical work centered around patient motivation toward prevention and tooth conservation. The scores for plaque and gingivitis were worse in the departments of oral surgery, dental materials, orthodontics and the basic science departments. Almost all departments and every class showed a few individuals with very poor oral hygiene. It is suggested that regular patient contact influences the personal attitude toward oral hygiene, and that professional activity and emphasis on different aspects of the curriculum may be reflected in the attitude of health professionals toward oral health.

Adult

Structural diversity and evolutionary constraints of oxidative phosphorylation.

The oxidative phosphorylation (OxPhos) system is central to metabolism. The more than 90 structural subunits are encoded by different chromosome categories (autosomal, X, and mtDNA). The system is envisioned as an invariant structure between cells and individuals. However, a comprehensive analysis of the 1,000 Genomes Project data reveals unexpected genetic intra-individual variability resulting from the heterozygosity of diploid autosomal genes, while diversity at the population level is generated by variability in mtDNA. We characterized the different levels of structural constriction at evolutionary and population levels for all OxPhos protein residues. To support this analysis, we developed ConScore, a conservation-based predictor of variant impact within OxPhos proteins (area under the receiver operating characteristic curve [ROC-AUC] = 0.97; area under the precision-recall curve [PR-AUC] = 0.94). Notably, for the nuclear-encoded subunits, we found mechanisms limiting individual variability as allelic imbalance or homozygosity bias. Integrating structural, functional, and genetic data, we highlight the significance of each OxPhos protein position, expanding insights into its role in speciation and disease.

Oxidative Phosphorylation

Architectural transcription factors collectively shape nuclear radial positioning of chromatin contacts.

The measurement of three-dimensional genome folding in the nucleus, mostly through Hi-C methods, is expressed as contact frequencies between genomic segments, without anchoring to physical axes of the spherical nucleus. Here, we mapped the chromatin contacts along nuclear radial axis and built radial score by factoring in contact frequencies. The chromatin high-order structures exhibit rich diversity along radial axis. Furthermore, the proximal trans contacts retrieved by radial score reveal conserved active/inactive chromatin segregation across intra- and interchromosomal interactions. Ablation of CTCF proteins disrupts chromatin loops with mild changes to chromatin radial positioning. By acutely perturbing multiple transcription factor (TF) occupancy, chromatin loop dissolutions are often accompanied by radial dissociations between two anchors. Our work provides a genome architecture reference map adhering to nuclear physical axis and suggests that multiple architectural TFs collectively shape nuclear positioning of chromatin and their contacts, with contacts serving as forces on chromatin positioning as well.

Chromatin

Screening for alcohol problems among the unemployed.

Of 2,996 welfare recipients applying for CETA benefits at the Milwaukee office of Jewish Vocational Service between 3/1/78-9/30/78, a 10% sample (N = 309) was screened for assessment of alcohol problems. After obtaining voluntary informed consent from participants (6% declined), trained interviewers individually administered a 16-item alcoholism At-Risk Questionnaire (ARQ) based on observations by NCA's Criteria Committee; a standard form of the 25-item Michigan Alcoholism Screening Test (MAST); and a 35-item interview structured around a selectively modified version of NCA's Criteria for the Diagnosis of Alcoholism (CRIT). Analyses of data suggested that our ARQ was of little value in discriminating between problem drinkers and other persons, although significantly correlated with MAST and CRIT scores. Using a conventional scoring of the MAST, 53.6% of the sample appeared to have significant alcohol problems, while our CRIT identified only 31.9% as problem drinkers. By combining the MAST + CRIT in a unique scoring system, a more conservative estimate of 36.57% problem drinkers, with an estimated error rate of 1.63% false negatives and 23.45% false positives, was determined. Further modification of MAST + CRIT scoring led to a revised estimate of 25.41% problem drinkers with estimated false-positive and false-negative rates of 7.55% and 6.5% respectively. Implications for research and plans for further modifications of screening procedures are discussed.

Adult

Beyond exons: Linking noncoding heritability and polygenicity across complex human traits and disorders.

The genetic architecture of complex traits spans a continuum of polygenicity, yet it remains unclear how differences in polygenicity relate to the functional localization of SNP heritability across the genome. We use a MiXeR-based framework to partition heritability across 74 functional annotations covering exonic, intronic, and intergenic regions for 34 complex traits and introduce a likelihood-based annotation contribution score that quantifies annotation-specific impact on heritability. Exons account for a minority of heritability, and their contribution decreases with increasing polygenicity, from an average of 22% in less-polygenic somatic diseases and biomarkers to 13% in highly polygenic psychiatric and cognitive phenotypes. Intergenic fractions show the opposite trend, whereas intronic fractions remain relatively stable. Analysis of the broader set of functional annotations also reveals systematic differences along the polygenicity axis: highly polygenic traits show stronger contributions from comparative genomics and variant-effect scores, whereas less-polygenic traits show stronger contributions from promoter, transcription, and chromatin annotations. Together, these results indicate that the functional partitioning of heritability systematically varies with polygenicity, shifting from gene-proximal regulatory architectures to architectures shaped by numerous dispersed regulatory effects.

MiXeR

Mutation rate heterogeneity biases variant effect prediction and reveals genuine mutational robustness.

Variant effect predictors (VEPs) are widely used to interpret the functional consequences of human genetic variation. Because most methods rely on sequence conservation, they implicitly treat conservation as evidence of functional constraint. However, substitution patterns across a phylogeny reflect not only selection but also differences in underlying mutation rates. Here, we show that this creates a systematic confounding: most VEPs capture mutation rate variation and misinterpret it as variation in functional importance. Widely used conservation metrics exhibit a related bias; in particular, phyloP scores correlate strongly with mutation rate even at putatively neutral sites. Consequently, variants at low-mutation-rate sites tend to be predicted as more damaging, and variants at highly mutable sites as more tolerated, than warranted by their true functional impact. We also identify a distinct biological signal in experimental measurements of mutational effects on protein stability: amino acid substitutions that are more likely to arise are, on average, less destabilizing than rarer substitutions. This provides empirical support for mutational robustness in the context of protein stability. However, this relationship is insufficient to explain the mutation-rate dependence observed in current VEP outputs. Together, our findings show that mutation rate heterogeneity systematically biases current variant effect prediction frameworks, highlight the need to model mutation probabilities explicitly in future VEPs, and reveal a genuine biological signal of mutational robustness.

conservation scores

Predicting student attrition in a baccalaureate curriculum.

A longitudinal study of enrollment data of the Universtiy of Wisconsin--Madison School of Nursing identified several trends in student attrition over a six-year period and formed the basis for the development of a model to predict student attrition. The model proposed to satisfy the need for more precise measures of probable student retention and attrition for counseling and planning purposes in the school's new curriculum. To arrive at a means for predicting student attrition, scores of random samples of continuing and dropout student on achievement, learning style, and psychologial variables obtained during the first three years of the new curriculum were examined using a discriminant analysis technique. A significant differentiation between continuing and dropout students (p less than .000) resulted and allowed a predictive function to be developed. Using a conservative method of identification with student scores transformed by this function, predictions of successful and nonsuccessful students were obtained.

Curriculum

Chromosome-level genome assembly and annotation of Pterygoplichthys pardalis.

Suckermouth catfishes, with their evolved powerful features, have become notorious invasive species, causing significant damage to aquatic ecosystems. However, the lack of high-quality genomes severely restricts research on this group within the field. In this study, we de novo assembled the chromosome-level genome assembly of Pterygoplichthys pardalis using multiple platforms of sequencing data, including Illumina short reads, Nanopore long reads, and Hi-C sequencing reads, resulting in a 1.51 Gb genome assembly. Multiple evaluations, including read mapping ratio (98.52%), transcript mapping ratio (99.61%), conserved BUSCO gene set (98.8%), and N50 score (49.47 Mb), indicated the high continuity and accuracy of the genome assembly we generated. Genome annotation found that 0.97 Gb of genome sequences are repetitive sequences, accounting for 64.47% of the genome assembly. Further, 23,859 protein-coding genes were successfully predicted, 92.92% of which could be annotated in functional databases. This high-quality genome assembly of P. pardalis provides a valuable resource for understanding the genetic underpinnings of P. pardalis's invasive success and offers critical data for future fisheries research and management.

Animals

[Determination of benzoic acid and its salts in food products of animal origin].

A modification of the method for determining benzoic acid and its salts by water-vapour distillation and extraction of the biological material conservant and its titrimetrical determination is proposed. The thin layer chromatography method (TLC) is used parallelly for the identification of the conservant. The advantage of the proposed method consists in that the same distiller is used for quantitative titrimetrical determination and for TLC assessment. The method possesses high sensitivity 1 gamma in TLC scoring and 0.1 mg in titrimetric determination. It is applied for the determination of conservants in various fish assortments, canned fish and roe.

Animals

Enrichment of root-associated Streptomyces strains in response to drought is driven by diverse functional traits and does not predict beneficial effects on plant growth.

The genus Streptomyces has consistently been found enriched in drought-stressed plant root microbiomes, yet the ecological basis and functional variation underlying this enrichment at the strain and isolate level remain unclear. Using two 16S rRNA sequencing methods with different levels of taxonomic resolution, we confirmed drought-associated enrichment (DE) of Streptomyces in field-grown sorghum roots and identified five closely related but distinct amplicon sequence variants (ASVs) belonging to the genus with variable drought enrichment patterns. From a culture collection of sorghum root endophytes, we selected 12 Streptomyces isolates representing these ASVs for phenotypic and genomic characterization. Whole-genome sequencing revealed substantial variation in gene content, even among closely related isolates, and exometabolomic profiling showed distinct metabolic responses to media supplemented with drought- versus well-watered root tissue. Traits linked to drought survival, including osmotic stress tolerance, siderophore production, and carbon utilization, varied widely among isolates and were not phylogenetically conserved. Using a broader panel of 48 Streptomyces, we demonstrate that DE scores, determined through mono-association experiments in gnotobiotic sorghum systems, showed high variability and lacked correlation with plant growth promotion. Pangenome-wide association identified orthogroups involved in osmolyte transport (e.g., proP) and membrane biosynthesis (e.g., fabG) as positively associated with DE, though most associations lacked phylogenetic signal. Collectively, these results demonstrate that Streptomyces DE is not a conserved genus-level trait but is instead strain-specific and functionally heterogeneous. Furthermore, DE in the root microbiome was shown not to predict beneficial effects on plant growth. This work underscores the need to resolve functional traits at the strain level and highlights the complexity of microbe-host-environment interactions under abiotic stress.

Streptomyces

ECLIPSE: exploring the dark proteome of ESKAPE pathogens through the sequence similarity network of the Protein Universe Atlas.

MOTIVATION: The accelerating crisis of antimicrobial resistance among the critical so-called ESKAPE pathogens demands the urgent identification of novel molecular targets. However, a substantial fraction of ESKAPE proteomes remains functionally uncharacterized, with many genes annotated as encoding hypothetical proteins. These protein sequences often lack significant similarity to known protein families when conventional homology-based annotation methods are used and thus remain "dark". This limits our ability to explore their roles in pathogenicity, and it is thus crucial to bridge this substantial gap in pathogen biology by developing new strategies to illuminate these "dark" regions of the ESKAPE pan-proteome. RESULTS: We introduce ECLIPSE (ESKAPE Connectome Linkage and Inference for Proteome Sequence Exploration), a network-based computational framework that systematically identifies and prioritizes functionally dark protein families in ESKAPE pan-proteomes. ECLIPSE embeds target ESKAPE pathogen proteomes within the global sequence similarity network of the Protein Universe Atlas. It detects connected components composed entirely of unannotated proteins, called the "dark proteome." As a case study, we applied ECLIPSE to a pan-proteome of 3 460 657 protein sequences from 635 strains of Pseudomonas aeruginosa (PA). ECLIPSE identified 120 985 proteins (4%) residing in completely dark connected components. Furthermore, we have performed a taxonomic diversity analysis using normalized Shannon indices to characterize each dark component by its enrichment in ESKAPE pathogens. The analysis utilized the evenness (E) value (see Methods 2.1), which distinguishes Pseudomonas-specific (target-specific) from ESKAPE-enriched dark components. We then developed the Dark Proteome Prioritization Score (DPPS), a composite multidimensional scoring framework (see Methods 2.5). It ranks these dark components by biological relevance across four orthogonal axes: (i) functional darkness, (ii) P. aeruginosa proportion in the Atlas, (iii) AMR-clade taxonomic restriction, and (iv) conservation across the 635 P. aeruginosa strains. This framework outputs a robust four-tier scoring system; the prioritized Tier I components were validated by weight sensitivity analysis and remained stable across 500 Monte Carlo weight perturbations. Structural characterization of one of the top-ranked ESKAPE-enriched dark components revealed that it belongs to the beta-barrel fold DUF1302 (PF06980) family, for which no experimentally solved three-dimensional structure exists in the PDB. The genomic context analysis indicates that it is co-localized with a LuxR-type transcriptional regulator. Collectively, ECLIPSE identifies evolutionarily conserved, structurally defined, and functionally dark proteins enriched across ESKAPE pathogens; these dark proteins can further be utilized as alternative antimicrobial targets for experimental characterization. AVAILABILITY AND IMPLEMENTATION: The source code and dataset are available for free at: Github: https://github.com/surabhilata/ECLIPSE.git, Zenodo: DOI: 10.5281/zenodo.21064323.

Proteome

Fast, accurate construction of multiple sequence alignments from protein language embeddings.

Multiple sequence alignment (MSA) is a foundational task in computational biology, underpinning protein structure prediction, evolutionary analysis, and domain annotation. Traditional MSA algorithms rely on pairwise amino acid substitution matrices derived from conserved protein families. While effective for aligning closely related sequences, these scoring schemes struggle in the low-identity "twilight zone." Here, we present a new approach for constructing MSAs leveraging amino acid embeddings generated by protein language models (PLMs), which capture rich evolutionary and contextual information from massive and diverse sequence datasets. We introduce a windowed reciprocal-weighted embedding similarity metric that is surprisingly effective in identifying corresponding amino acids across sequences. Building on this metric, we develop ARIES (Alignment via RecIprocal Embedding Similarity), an algorithm that constructs a PLM-generated template embedding and aligns each sequence to this template via dynamic time warping in order to build a global MSA. Across diverse benchmark datasets, ARIES achieves higher accuracies than existing state-of-the-art approaches, especially in low-identity regimes where traditional methods degrade, while scaling almost linearly with the number of sequences to be aligned. Together, these results provide the first large-scale demonstration of the power of PLMs for accurate and scalable MSA construction across protein families of varying sizes and levels of similarity, highlighting the potential of PLMs to transform comparative sequence analysis.

Deep Learning

NCBoost v2: a classifier for non-coding single-nucleotide variants in Mendelian diseases.

MOTIVATION: The current diagnostic rate of rare diseases through whole-genome sequencing has stabilized at around 30% on average, highlighting the need for improved computational scores to identify pathogenic variants. In 2019, we developed NCBoost, a supervised-learning approach that mined a comprehensive set of sequence constraint features and proved particularly well suited to identifying high-effect pathogenic non-coding variants in genetic diseases. Since its first release, the substantial increase in the number of variants available for training, as well as the enhanced capacity to detect purifying selection signals from large-scale genome sequencing projects, motivated an update of NCBoost. RESULTS: We implemented NCBoost v2, a pathogenicity score for non-coding single-nucleotide variants, trained on the largest set of curated pathogenic variants in monogenic Mendelian diseases available to date. It leverages conservation features computed from recent large-scale genomic consortia such as Zoonomia and gnomAD, and incorporates recent splice-altering predictive scores. NCBoost v2 outperformed alternative state-of-the-art methods in a variety of scenarii, providing more consistent scores across non-coding genomic regions and fine-tuning the scoring of pathogenic splice-altering variants in Mendelian disease genes. AVAILABILITY AND IMPLEMENTATION: NCBoost v2 software is implemented in Python 3.10 and is freely available under the GNU General Public License Version 3 at https://doi.org/10.5281/zenodo.16029049 and https://github.com/RausellLab/NCBoost-2, together with precomputed scores for the human genome assembly GRCh38.

Polymorphism, Single Nucleotide

Molecular Cloning, Recombinant Expression, and In Silico Structural Analysis of Cu/Zn-Superoxide Dismutase from Trachyspermum ammi.

Superoxide dismutase (SOD) is an essential antioxidant metalloenzyme that is critical for the cellular defense against oxidative damage, as it scavenges superoxide radicals and maintains the redox status. Cytosolic Cu/Zn-SOD is particularly important in the regulation of oxidative stress among different isoforms in higher plants. While Cu/Zn-SODs from several plant species have been characterized, molecular information is limited for Trachyspermum ammi, a medicinally important member of a family Apiaceae with antioxidant potential.In the present study, an integrated molecular and in silico approach has been taken to clone and analyze a Cu/Zn type SOD gene from T. ammi to get insight into its structural and evolutionary characteristics. PCR amplification yielded an open reading frame of 456 bp encoding a protein of 152 amino acids. Sequence analysis showed that plant Cu/Zn-SODs, especially those from Daucus carota, were highly similar to one another (about 90-95%).Multiple sequence alignment confirmed the presence of conserved catalytic motifs and metal-binding histidine residues, both of which are crucial for enzymatic function. Physicochemical analysis predicted the protein to be stable, hydrophilic and compatible with cytosolic localization. The analysis of secondary structure indicated a predominance of β-strands, consistent with the conserved β-barrel architecture of plant Cu/Zn-SODs.The three-dimensional structure was built by homology modeling using a closely related plant Cu/Zn-SOD template with high sequence identity. Structural validation demonstrated an acceptable stereochemical quality with 86.3% residues in the favored region of Ramachandran plot, satisfactory ERRAT and Verify3D scores, and a low RMSD value of 0.104 Å on structural superimposition. Phylogenetic analysis placed the enzyme in the Apiaceae lineage, suggesting evolutionary conservation among related plant species. In conclusion, this study presents the first molecular and structural characterization of Cu/Zn-SOD from T. ammi and confirms the existence of a conserved structural framework typical of plant Cu/Zn-SODs. These results provide a basis for further studies concerning recombinant expression, enzymatic validation and potential relevance in antioxidant and plant stress biology.

Cloning, Molecular