Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,117 records · Page 62Linked to original sources

A simple method for statistical analysis of intensity differences in microarray-derived gene expression data.

BACKGROUND: Microarray experiments offer a potent solution to the problem of making and comparing large numbers of gene expression measurements either in different cell types or in the same cell type under different conditions. Inferences about the biological relevance of observed changes in expression depend on the statistical significance of the changes. In lieu of many replicates with which to determine accurate intensity means and variances, reliable estimates of statistical significance remain problematic. Without such estimates, overly conservative choices for significance must be enforced. RESULTS: A simple statistical method for estimating variances from microarray control data which does not require multiple replicates is presented. Comparison of datasets from two commercial entities using this difference-averaging method demonstrates that the standard deviation of the signal scales at a level intermediate between the signal intensity and its square root. Application of the method to a dataset related to the beta-catenin pathway yields a larger number of biologically reasonable genes whose expression is altered than the ratio method. CONCLUSIONS: The difference-averaging method enables determination of variances as a function of signal intensities by averaging over the entire dataset. The method also provides a platform-independent view of important statistical properties of microarray data.

Algorithms↗

Measurement of fractionated plasma metanephrines for exclusion of pheochromocytoma: Can specificity be improved by adjustment for age?

BACKGROUND: Biochemical testing for pheochromocytoma by measurement of fractionated plasma metanephrines is limited by false positive rates of up to 18% in people without known genetic predisposition to the disease. The plasma normetanephrine fraction is responsible for most false positives and plasma normetanephrine increases with age. The objective of this study was to determine if we could improve the specificity of fractionated plasma measurements, by statistically adjusting for age. METHODS: An age-adjusted metanephrine score was derived using logistic regression from 343 subjects (including 33 people with pheochromocytoma) who underwent fractionated plasma metanephrine measurements as part of investigations for suspected pheochromocytoma at Mayo Clinic Rochester (derivation set). The performance of the age-adjusted score was validated in a dataset of 158 subjects (including patients 23 with pheochromocytoma) that underwent measurements of fractionated plasma metanephrines at Mayo Clinic the following year (validation dataset). None of the participants in the validation dataset had known genetic predisposition to pheochromocytoma. RESULTS: The sensitivity of the age-adjusted metanephrine score was the same as that of traditional interpretation of fractionated plasma metanephrine measurements, yielding a sensitivity of 100% (23/23, 95% confidence interval [CI] 85.7%, 100%). However, the false positive rate with traditional interpretation of fractionated plasma metanephrine measurements was 16.3% (22/135, 95% CI, 11.0%, 23.4%) and that of the age-adjusted score was significantly lower at 3.0% (4/135, 95% CI, 1.2%, 7.4%) (p < 0.001 using McNemar's test). CONCLUSION: An adjustment for age in the interpretation of results of fractionated plasma metanephrines may significantly decrease false positives when using this test to exclude sporadic pheochromocytoma. Such improvements in false positive rate may result in savings of expenditures related to confirmatory imaging.

Journal Article↗

The use of bootstrap methods for analysing Health-Related Quality of Life outcomes (particularly the SF-36).

Health-Related Quality of Life (HRQoL) measures are becoming increasingly used in clinical trials as primary outcome measures. Investigators are now asking statisticians for advice on how to analyse studies that have used HRQoL outcomes.HRQoL outcomes, like the SF-36, are usually measured on an ordinal scale. However, most investigators assume that there exists an underlying continuous latent variable that measures HRQoL, and that the actual measured outcomes (the ordered categories), reflect contiguous intervals along this continuum. The ordinal scaling of HRQoL measures means they tend to generate data that have discrete, bounded and skewed distributions. Thus, standard methods of analysis such as the t-test and linear regression that assume Normality and constant variance may not be appropriate. For this reason, conventional statistical advice would suggest that non-parametric methods be used to analyse HRQoL data. The bootstrap is one such computer intensive non-parametric method for analysing data. We used the bootstrap for hypothesis testing and the estimation of standard errors and confidence intervals for parameters, in four datasets (which illustrate the different aspects of study design). We then compared and contrasted the bootstrap with standard methods of analysing HRQoL outcomes. The standard methods included t-tests, linear regression, summary measures and General Linear Models.Overall, in the datasets we studied, using the SF-36 outcome, bootstrap methods produce results similar to conventional statistical methods. This is likely because the t-test and linear regression are robust to the violations of assumptions that HRQoL data are likely to cause (i.e. non-Normality). While particular to our datasets, these findings are likely to generalise to other HRQoL outcomes, which have discrete, bounded and skewed distributions. Future research with other HRQoL outcome measures, interventions and populations, is required to confirm this conclusion.

Arthritis, Rheumatoid↗

Distinguishing HIV-1 drug resistance, accessory, and viral fitness mutations using conditional selection pressure analysis of treated versus untreated patient samples.

BACKGROUND: HIV can evolve drug resistance rapidly in response to new drug treatments, often through a combination of multiple mutations 123. It would be useful to develop automated analyses of HIV sequence polymorphism that are able to predict drug resistance mutations, and to distinguish different types of functional roles among such mutations, for example, those that directly cause drug resistance, versus those that play an accessory role. Detecting functional interactions between mutations is essential for this classification. We have adapted a well-known measure of evolutionary selection pressure (Ka/Ks) and developed a conditional Ka/Ks approach to detect important interactions. RESULTS: We have applied this analysis to four independent HIV protease sequencing datasets: 50,000 clinical samples sequenced by Specialty Laboratories, Inc.; 1800 samples from patients treated with protease inhibitors; 2600 samples from untreated patients; 400 samples from untreated African patients. We have identified 428 mutation interactions in Specialty dataset with statistical significance and we were able to distinguish primary vs. accessory mutations for many well-studied examples. Amino acid interactions identified by conditional Ka/Ks matched 80 of 92 pair wise interactions found by a completely independent study of HIV protease (p-value for this match is significant: 10-70). Furthermore, Ka/Ks selection pressure results were highly reproducible among these independent datasets, both qualitatively and quantitatively, suggesting that they are detecting real drug-resistance and viral fitness mutations in the wild HIV-1 population. CONCLUSION: Conditional Ka/Ks analysis can detect mutation interactions and distinguish primary vs. accessory mutations in HIV-1. Ka/Ks analysis of treated vs. untreated patient data can distinguish drug-resistance vs. viral fitness mutations. Verification of these results would require longitudinal studies. The result provides a valuable resource for AIDS research and will be available for open access upon publication at http://www.bioinformatics.ucla.edu/HIV.

Journal Article↗

Tailor-made composite functions as tools in model choice: the case of sigmoidal vs bi-linear growth profiles.

BACKGROUND: Roots are the classical model system to study the organization and dynamics of organ growth zones. Profiles of the velocity of root elements relative to the apex have generally been considered to be sigmoidal. However, recent high-resolution measurements have yielded bi-linear profiles, suggesting that sigmoidal profiles may be artifacts caused by insufficient spatio-temporal resolution. The decision whether an empirical velocity profile follows a sigmoidal or bi-linear distribution has consequences for the interpretation of the underlying biological processes. However, distinguishing between sigmoidal and bi-linear curves is notoriously problematic. A mathematical function that can describe both types of curve equally well would allow them to be distinguished by automated curve-fitting. RESULTS: On the basis of the mathematical requirements defined, we created a composite function and tested it by fitting it to sigmoidal and bi-linear models with different noise levels (Monte-Carlo datasets) and to three experimental datasets from roots of Gypsophila elegans, Aurinia saxatilis, and Arabidopsis thaliana. Fits of the function proved robust with respect to noise and yielded statistically sound results if care was taken to identify reasonable initial coefficient values to start the automated fitting procedure. Descriptions of experimental datasets were significantly better than those provided by the Richards function, the most flexible of the classical growth equations, even in cases in which the data followed a smooth sigmoidal distribution. CONCLUSION: Fits of the composite function introduced here provide an independent criterion for distinguishing sigmoidal and bi-linear growth profiles, but without forcing a dichotomous decision, as intermediate solutions are possible. Our function thus facilitates an unbiased, multiple-working hypothesis approach. While our discussion focusses on kinematic growth analysis, this and similar tailor-made functions will be useful tools wherever models of steadily or abruptly changing dependencies between empirical parameters are to be compared.

Journal Article↗

An expanded genome-scale model of Escherichia coli K-12 (iJR904 GSM/GPR).

BACKGROUND: Diverse datasets, including genomic, transcriptomic, proteomic and metabolomic data, are becoming readily available for specific organisms. There is currently a need to integrate these datasets within an in silico modeling framework. Constraint-based models of Escherichia coli K-12 MG1655 have been developed and used to study the bacterium's metabolism and phenotypic behavior. The most comprehensive E. coli model to date (E. coli iJE660a GSM) accounts for 660 genes and includes 627 unique biochemical reactions. RESULTS: An expanded genome-scale metabolic model of E. coli (iJR904 GSM/GPR) has been reconstructed which includes 904 genes and 931 unique biochemical reactions. The reactions in the expanded model are both elementally and charge balanced. Network gap analysis led to putative assignments for 55 open reading frames (ORFs). Gene to protein to reaction associations (GPR) are now directly included in the model. Comparisons between predictions made by iJR904 and iJE660a models show that they are generally similar but differ under certain circumstances. Analysis of genome-scale proton balancing shows how the flux of protons into and out of the medium is important for maximizing cellular growth. CONCLUSIONS: E. coli iJR904 has improved capabilities over iJE660a. iJR904 is a more complete and chemically accurate description of E. coli metabolism than iJE660a. Perhaps most importantly, iJR904 can be used for analyzing and integrating the diverse datasets. iJR904 will help to outline the genotype-phenotype relationship for E. coli K-12, as it can account for genomic, transcriptomic, proteomic and fluxomic data simultaneously.

Citric Acid Cycle↗

A first-draft human protein-interaction map.

BACKGROUND: Protein-interaction maps are powerful tools for suggesting the cellular functions of genes. Although large-scale protein-interaction maps have been generated for several invertebrate species, projects of a similar scale have not yet been described for any mammal. Because many physical interactions are conserved between species, it should be possible to infer information about human protein interactions (and hence protein function) using model organism protein-interaction datasets. RESULTS: Here we describe a network of over 70,000 predicted physical interactions between around 6,200 human proteins generated using the data from lower eukaryotic protein-interaction maps. The physiological relevance of this network is supported by its ability to preferentially connect human proteins that share the same functional annotations, and we show how the network can be used to successfully predict the functions of human proteins. We find that combining interaction datasets from a single organism (but generated using independent assays) and combining interaction datasets from two organisms (but generated using the same assay) are both very effective ways of further improving the accuracy of protein-interaction maps. CONCLUSIONS: The complete network predicts interactions for a third of human genes, including 448 human disease genes and 1,482 genes of unknown function, and so provides a rich framework for biomedical research.

Databases, Protein↗

The expression signature of in vitro senescence resembles mouse but not human aging.

BACKGROUND: The biological mechanisms that underlie aging have not yet been fully identified. Senescence, a phenomenon occurring in vitro, limits the number of cell divisions in mammalian cell cultures and has been suggested to contribute to aging. RESULTS: We investigated whether the changes in gene expression that occur during mammalian aging and induction of cellular senescence are similar. We compared changes of gene expression in seven microarray datasets from aging human, mouse and rat, as well as four microarray datasets from senescent cells of man and mouse. The datasets were publicly available or obtained from other laboratories. Correlation measures were used to establish similarities of the expression profiles and gene ontology analyses to identify functional groups of genes that are co-regulated. Robust similarities were established between aging in different species and tissues, indicating that there is an aging transcriptome. Although some cross-species comparisons displayed high correlation, intra-species similarities were more reliable. Similarly, a senescence transcriptome was demonstrated that is conserved across cell types. A similarity between the expression signatures of cellular senescence and aging could be established in mouse, but not in human. CONCLUSION: Our study is the first to use microarray data from several studies and laboratories for dissection of a complex biological phenotype. We demonstrate the presence of a mammalian aging transcriptome, and discuss why similarity between cellular senescence and aging is apparent in aging mice only.

Aging↗

LungGENIE: the lung gene-expression and network imputation engine.

BACKGROUND: Few cohorts have study populations large enough to conduct molecular analysis of ex vivo lung tissue for genomic analyses. Transcriptome imputation is a non-invasive alternative with many potential applications. We present a novel transcriptome-imputation method called the Lung Gene Expression and Network Imputation Engine (LungGENIE) that uses principal components from blood gene-expression levels in a linear regression model to predict lung tissue-specific gene-expression. METHODS: We use paired blood and lung RNA sequencing data from the Genotype-Tissue Expression (GTEx) project to train LungGENIE models. We replicate model performance in a unique dataset, where we generated RNA sequencing data from paired lung and blood samples available through the SUNY Upstate Biorepository (SUBR). We further demonstrate proof-of-concept application of LungGENIE models in an independent blood RNA sequencing data from the Genetic Epidemiology of COPD (COPDGene) study. RESULTS: We show that LungGENIE prediction accuracies have higher correlation to measured lung tissue expression compared to existing cis-expression quantitative trait loci-based methods (median Pearson's r&#x2009;=&#x2009;0.25, IQR 0.19-0.32), with close to half of the reliably predicted transcripts being replicated in the testing dataset. Finally, we demonstrate significant correlation of differential expression results in chronic obstructive pulmonary disease (COPD) from imputed lung tissue gene-expression and differential expression results experimentally determined from lung tissue. CONCLUSION: Our results demonstrate that LungGENIE provides complementary results to existing expression quantitative trait loci-based methods and outperforms direct blood to lung results across internal cross-validation, external replication, and proof-of-concept in an independent dataset. Taken together, we establish LungGENIE as a tool with many potential applications in the study of lung diseases.

Humans↗

scFANCL: Dual contrastive learning with false-negative correction at cell level for single-cell RNA-seq clustering.

BACKGROUND: Single-cell RNA sequencing (scRNA-seq) enables cellular characterization at single-cell resolution. However, its high dimensionality, sparsity, and noise make clustering challenging. Approaches utilizing contrastive learning and data augmentation have been introduced to improve representation quality for scRNA-seq clustering. In particular, dual contrastive frameworks combining instance- and cluster-level objectives can capture both cell-cell similarities and inter-cluster variations. However, existing dual contrastive frameworks focus primarily on discrete cluster boundaries, neglecting the biological continuity inherent in scRNA-seq data. METHODS: We propose scFANCL, a dual contrastive framework designed to capture biological continuity in scRNA data. Rather than treating all non-augmented samples as negatives, scFANCL applies a cosine-similarity-based threshold to exclude cells of the same type from the negative pool, preserving continuous transcriptional relationships among them while maintaining inter-cluster separation. RESULTS: Extensive experiments across seven publicly available scRNA-seq datasets demonstrated that scFANCL achieves competitive clustering performance compared with existing baseline methods, consistently yielding high ARI and NMI scores across datasets of varying size and complexity. Ablation studies further confirmed the contribution of the false negative filtering component, showing measurable improvements over variants without filtering. Downstream analyses further suggest that the learned embeddings may reflect biologically meaningful transcriptional transitions, including continuous differentiation trajectories within related cell types. The source code is available at https://github.com/mjuailab/scFANCL . CONCLUSIONS: scFANCL addresses a key limitation of conventional contrastive learning by applying a cosine-similarity-based threshold to exclude cells of the same type from the negative pool, thereby preserving biological continuity within cell types while maintaining inter-cluster separation. Evaluations across seven benchmark scRNA-seq datasets demonstrate competitive clustering performance, with learned embeddings capturing biologically meaningful transcriptional structure and characteristics of rare cell populations.

Clustering Algorithms↗

Exploring shotgun metagenomic data to detect microeukaryotic pathogens in wildlife.

BACKGROUND: Microeukaryotic parasites of the intestinal tract are an understudied group of organisms that infect humans and many other animals. Targeted sequencing methods focused on individual loci are usually employed for detection of these parasites, making comprehensive studies of microeukaryotic parasite diversity within hosts or other systems difficult. Exploratory approaches such as shotgun metagenomic sequencing to survey the diversity of microeukaryotic parasites in new and existing datasets are not well developed. RESULTS: Utilizing existing datasets from 12 goose fecal samples, we explored some of the benefits and challenges of using shotgun metagenome sequencing to detect microeukaryotic parasites. We demonstrated the importance of careful curation of read classification data to avoid erroneously linking pathogens to hosts or environments as unsupported classifications were common in the data and varied widely depending on analysis parameters. However, we were able to establish strong support for the presence of sequences of Eimeria and Enterocytozoon bieneusi. In addition, examination of trichomonad reads indicated that parasite reads mapping to human pathogens unlikely to colonize geese may in fact represent cryptic microeukaryotic species that are not included in existing curated databases opening new potential avenues of study. CONCLUSIONS: Taken together these findings support the idea that exploring microeukaryotic parasite diversity within shotgun metagenomic datasets can be beneficial to our understanding of the presence and diversity of these organisms in wildlife hosts.

Animals↗

Shared genetic architecture between ADHD and intelligence varies across ADHD subtypes.

BACKGROUND: Attention-deficit/hyperactivity disorder (ADHD) is a heterogeneous neurodevelopmental condition frequently accompanied by cognitive difficulties. Although previous genetic studies have demonstrated substantial overlap between ADHD and intelligence, most have treated ADHD as a single phenotype. However, whether this shared genetic architecture differs across ADHD subtypes remains unclear. METHODS: We conducted a genome-wide cross-trait analysis integrating large-scale genome-wide association study (GWAS) datasets of overall ADHD, its subtypes-childhood ADHD, persistent ADHD, and late-diagnosed ADHD-and intelligence (total N&#x2009;>&#x2009;300,000). Genome-wide genetic correlations, polygenic overlap, local genetic correlations, and variant-level associations between ADHD phenotypes and intelligence were evaluated to characterize their shared genetic architecture. Shared variants were identified through cross-trait enrichment analyses and subsequently mapped to genes for functional annotation and gene-set enrichment. Bidirectional associations were evaluated using two-sample Mendelian randomization with sensitivity analyses. Additional GWAS datasets were used to validate the robustness of shared loci by assessing the consistency of effect directions. RESULTS: All ADHD phenotypes showed significant negative genetic correlations with intelligence (rg ranging from -0.3442 to -0.4205). Despite these modest genome-wide correlations, cross-trait analyses revealed substantial genetic overlap, including polygenic overlap, local genetic correlations, and variant-level associations. We identified 184 loci jointly associated with ADHD traits and intelligence, including 64 novel loci, whereas no shared loci were detected for persistent ADHD under the current analysis. Functional annotation revealed biologically distinct enrichment patterns across subtypes: childhood ADHD loci were linked to early neurodevelopmental processes, while late-diagnosed ADHD loci were enriched in synapse-related and neuronal signaling pathways. Mendelian randomization analyses suggested bidirectional associations, with stronger evidence supporting a directional association from intelligence to ADHD risk. Furthermore, these shared loci showed largely consistent effect directions across additional GWAS datasets, providing support for the robustness of the findings. CONCLUSIONS: The shared genetic architecture between ADHD and intelligence varies across ADHD subtypes, highlighting distinct biological pathways underlying cognitive heterogeneity in ADHD. These findings suggest that the relationship between ADHD liability and general cognitive ability is not uniform across ADHD subtypes and may inform future research on risk stratification and early identification in child and adolescent psychiatry.

Humans↗

N6-methyladenine identification using deep learning and discriminative feature integration.

N6-methyladenine (6&#xa0;mA) is a pivotal DNA modification that plays a crucial role in epigenetic regulation, gene expression, and various biological processes. With advancements in sequencing technologies and computational biology, there is an increasing focus on developing accurate methods for 6&#xa0;mA site identification to enhance early detection and understand its biological significance. Despite the rapid progress of machine learning in bioinformatics, accurately detecting 6&#xa0;mA sites remains a challenge due to the limited generalizability and efficiency of existing approaches. In this study, we present Deep-N6mA, a novel Deep Neural Network (DNN) model incorporating optimal hybrid features for precise 6&#xa0;mA site identification. The proposed framework captures complex patterns from DNA sequences through a comprehensive feature extraction process, leveraging k-mer, Dinucleotide-based Cross Covariance (DCC), Trinucleotide-based Auto Covariance (TAC), Pseudo Single Nucleotide Composition (PseSNC), Pseudo Dinucleotide Composition (PseDNC), and Pseudo Trinucleotide Composition (PseTNC). To optimize computational efficiency and eliminate irrelevant or noisy features, an unsupervised Principal Component Analysis (PCA) algorithm is employed, ensuring the selection of the most informative features. A multilayer DNN serves as the classification algorithm to identify N6-methyladenine sites accurately. The robustness and generalizability of Deep-N6mA were rigorously validated using fivefold cross-validation on two benchmark datasets. Experimental results reveal that Deep-N6mA achieves an average accuracy of 97.70% on the F. vesca dataset and 95.75% on the R. chinensis dataset, outperforming existing methods by 4.12% and 4.55%, respectively. These findings underscore the effectiveness of Deep-N6mA as a reliable tool for early 6&#xa0;mA site detection, contributing to epigenetic research and advancing the field of computational biology.

Deep Learning↗

Sugar-sweetened beverage consumption and incident depression: an exploratory multi-omics analysis of candidate biological mediators.

BACKGROUND: Depression is a leading cause of mental and physical disability globally, with its onset and progression influenced by a complex interplay of dietary, psychological, and biological factors. Recent research suggests a link between sugar-sweetened beverage (SSB) consumption and depression risk, although the potential biological pathways underlying this association remain poorly understood. METHODS: This study utilized data from 192,045 participants in the UK Biobank to examine the prospective association between SSB consumption and incident depression using Cox proportional hazards models. SSBs were defined as the sum of five beverage categories assessed via the Oxford WebQ 24-hour dietary recall. Directional consistency of the association was further examined across three external supporting datasets encompassing diverse populations: NHANES, YRBSS, and the Lianyungang Municipal School Health and Risk Factor Surveillance Study Dataset. We further investigated whether proteins, metabolites, inflammatory markers, and brain imaging phenotypes may serve as candidate mediators statistically consistent with mediation of the SSB-depression association. RESULTS: High SSB consumption was associated with an 18% higher risk of incident depression compared with non-consumers (HR&#x2009;=&#x2009;1.18; 95% CI: 1.11-1.25), with consistent directional associations observed across external supporting datasets. A plasma proteomic signature comprising 229 proteins was constructed using elastic net regularization and was associated with an increased risk of incident depression. Exploratory mediation analyses identified 72 proteins, 36 metabolites, and 5 inflammatory markers as candidate mediators, with IL1RN showing the strongest protein-level candidate mediating effect (9.6%), and Unsaturation and neutrophil count showing the strongest metabolite- and inflammatory marker-level effects, respectively. CONCLUSIONS: This study provides preliminary evidence that proteins, metabolites, and inflammatory markers may serve as candidate mediators statistically consistent with mediation of the association between SSB consumption and incident depression. These findings are exploratory and hypothesis-generating, and future experimental studies are needed to validate these candidate pathways and assess their potential as targets for dietary interventions in depression prevention.

Humans↗

Beyond genes: EpiSwitch&#xae; and Orion platform-powered 3D genome architecture biomarkers reveal shared biology across ME/CFS, long COVID, PTSD, rheumatoid arthritis, and multiple sclerosis.

BACKGROUND: Myalgic encephalomyelitis/chronic fatigue syndrome (ME/CFS), Long COVID (LC19), post-traumatic stress disorder (PTSD), rheumatoid arthritis (RA), and multiple sclerosis (MS) are clinically distinct disorders that share substantial symptom overlap, including persistent fatigue, cognitive impairment, autonomic dysfunction, and immune dysregulation. Although these conditions differ in diagnosis and clinical presentation, their underlying biological mechanisms remain poorly understood and may involve convergent regulatory pathways. METHODS: The EpiSwitch&#xae; 3D genomics platform and Orion knowledgebase were used to integrate chromosome conformation signatures with genome-wide association study (GWAS)-derived datasets across ME/CFS, LC19, PTSD, RA, and MS. Three-dimensional genomic anchors were mapped to coding genes and analysed using STRING protein-protein interaction networks and Cytoscape-based systems biology approaches. Disease-specific anchor datasets were generated and compared at both gene and network levels to identify shared biological processes and regulatory mechanisms. RESULTS: Analysis of the ME/CFS dataset identified 552 unique 3D genomic anchors mapped to 567 genes, with analogous disease-specific anchor sets generated for LC19, PTSD, RA, and MS. Direct overlap between disease-associated genes was limited; however, higher-order network analyses revealed substantial interconnectivity and convergence across conditions. Shared biological pathways included immune and cytokine signalling, interferon responses, mitochondrial function, metabolic regulation, and neuroendocrine processes. Highly connected hub genes included immune regulatory nodes such as LAG3 and components of the mTOR signalling pathway, implicating T-cell exhaustion, chronic immune activation, and immunometabolic dysregulation as common mechanisms underlying these disorders. CONCLUSIONS: These findings support a systems-level model in which clinically overlapping fatigue-associated syndromes arise from perturbations of interconnected regulatory networks rather than discrete disease-specific pathways. Despite limited genetic overlap, substantial convergence at the network level suggests shared biological architecture across ME/CFS, LC19, PTSD, RA, and MS. The identification of common regulatory pathways provides a mechanistic framework for the development of cross-disease diagnostic and therapeutic strategies. By capturing dynamic regulatory states, 3D genomic biomarkers offer significant potential for objective blood-based diagnostics, patient stratification, and the identification of shared therapeutic targets across complex chronic disorders. These findings support the application of precision medicine approaches and may accelerate the development of novel interventions for fatigue-associated multisystem diseases.

Humans↗

A polymorphism in the cystatin C gene is a novel risk factor for late-onset Alzheimer's disease.

OBJECTIVE: To investigate whether or not a coding polymorphism in the cystatin C gene (CST3) contributes risk for AD. DESIGN: A case-control genetic association study of a Caucasian dataset of 309 clinic- and community-based cases and 134 community-based controls. RESULTS: The authors find a signficant interaction between the GG genotype of CST3 and age/age of onset on risk for AD, such that in the over-80 age group the GG genotype contributes two-fold increased risk for the disease. The authors also see a trend toward interaction between APOE epsilon4-carrying genotype and age/age of onset in this dataset, but in the case of APOE the risk decreases with age. Analysis of only the community-based cases versus controls reveals a significant three-way interaction between APOE, CST3 and age/age of onset. CONCLUSION: The reduced or absent risk for AD conferred by APOE in older populations has been well reported in the literature, prompting the suggestion that additional genetic risk factors confer risk for later-onset AD. In the author's dataset the opposite effects of APOE and CST3 genotype on risk for AD with increasing age suggest that CST3 is one of the risk factors for later-onset AD. Although the functional significance of this coding polymorphism has not yet been reported, several hypotheses can be proposed as to how variation in an amyloidogenic cysteine protease inhibitor may have pathologic consequences for AD.

Aged↗

Relaxed phylogenetics and dating with confidence.

In phylogenetics, the unrooted model of phylogeny and the strict molecular clock model are two extremes of a continuum. Despite their dominance in phylogenetic inference, it is evident that both are biologically unrealistic and that the real evolutionary process lies between these two extremes. Fortunately, intermediate models employing relaxed molecular clocks have been described. These models open the gate to a new field of "relaxed phylogenetics." Here we introduce a new approach to performing relaxed phylogenetic analysis. We describe how it can be used to estimate phylogenies and divergence times in the face of uncertainty in evolutionary rates and calibration times. Our approach also provides a means for measuring the clocklikeness of datasets and comparing this measure between different genes and phylogenies. We find no significant rate autocorrelation among branches in three large datasets, suggesting that autocorrelated models are not necessarily suitable for these data. In addition, we place these datasets on the continuum of clocklikeness between a strict molecular clock and the alternative unrooted extreme. Finally, we present analyses of 102 bacterial, 106 yeast, 61 plant, 99 metazoan, and 500 primate alignments. From these we conclude that our method is phylogenetically more accurate and precise than the traditional unrooted model while adding the ability to infer a timescale to evolution.

Animals↗

Evolutionary and physiological importance of hub proteins.

It has been claimed that proteins with more interaction partners (hubs) are both physiologically more important (i.e., less dispensable) and, owing to an assumed high density of binding sites, slow evolving. Not all analyses, however, support these results, probably because of biased and less-than reliable global protein interaction data. Here we provide the first examination of these issues using a comprehensive literature-curated dataset of well-substantiated protein interactions in Saccharomyces cerevisiae. Whereas use of less reliable yeast two-hybrid data alone can reject the possibility that local connectivity correlates with measures of dispensability, in higher quality datasets a relatively robust correlation is observed. In contrast, local connectivity does not correlate with the rate of protein evolution even in reliable datasets. This perhaps surprising lack of correlation with evolutionary rate appears in part to arise from the fact that hub proteins do not have a higher density of residues associated with binding. However, hub proteins do have at least one other set of unusual features, namely rapid turnover and regulation, as manifest in high mRNA decay rates and a large number of phosphorylation sites. This, we suggest, is an adaptation to minimize unwanted activation of pathways that might be mediated by adventitious binding to hubs, were they to actively persist longer than required at any given time point. We conclude that hub proteins are more important for cellular growth rate and under tight regulation but are not slow evolving.

Biological Evolution↗