Search PubMedSearch

SEARCH · Search PubMed

Results for “big data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Big data in multiple sclerosis.

PURPOSE OF REVIEW: This review summarizes recent key advancements in multiple sclerosis (MS) achieved through the utilization of big data from diverse sources and advanced analytical techniques. RECENT FINDINGS: Real-world evidence (RWE) derived from MS big data has significantly enhanced treatment strategies, redefined the concept of disease progression, refined prognostic models, and facilitated personalized medicine. RWE has highlighted the long-term benefits of early intensive treatment compared to escalation strategies, the unfavorable risk profile associated with treatment de-escalation and the importance of managing treatments during pregnancy. Additionally, it has revealed similarities and differences in the effectiveness and safety of specific high-efficacy therapies, as well as key predictors for switching treatments. RWE has also emphasized the central role of progression independent of relapse activity as a significant driver of disability and predictor of unfavorable long-term outcomes in both adult and pediatric onset MS. A data-driven approach utilizing artificial intelligence and big data has established a comprehensive framework for understanding the disease's evolution. Multimodal big data frameworks - encompassing clinical data, MRI, genomics, biomarkers, and app-based metrics - have demonstrated their ability to enhance diagnostic performance and risk stratification in MS. SUMMARY: Big data approaches are transforming MS research and clinical practice by providing stronger RWE to guide therapeutic decision-making, refining models of disease progression, and developing more precise prognostic tools.

Humans

Big data and psychiatry: advances, constraints and future directions.

Early work in psychiatry research, often involving single sites, small samples, and limited variables, has shifted to contemporary research involving multiple sites, large samples, and many variables. Such research raises important questions, including concerns about data quality and methodological rigor, uncertainty about its key lessons, issues regarding clinical relevance, and questions about how to optimize future advances. Here we consider these questions and concerns against the context of big data work on community and register-based surveys, cohort and biobank studies, electronic health records, digital phenotyping, brain imaging, genomics and other -omics, and randomized controlled trials. The development of large datasets allowing well-powered analyses is a major milestone, but sample size alone does not guarantee more precise estimates, and ongoing attention to the quality and rigor of big data collation and analysis is needed. Big data research has fostered trans-disciplinarity and given insights into mechanisms underlying psychiatric disorders, but also emphasizes the intricacy, heterogeneity and variability of such mechanisms, and the importance of triangulating between large-scale and small-scale research. The complexity of psychiatric phenotypes and psychobiological mechanisms contributes to the difficulty in bridging from big data to clinical application; big data research reinforces the importance of holding our diagnoses of psychiatric disorders lightly and providing explanations of these conditions humbly; and future work needs to be more attentive to clinical issues. There is enormous scope for further building databases relevant to psychiatry, but advances in conceptual models and asking the right questions are equally valuable. The full impact of big data, including artificial intelligence analyses, remains to be seen, but overenthusiastic support should be tempered by a better understanding of its strengths and limitations. At its best, such work will contribute in an iterative and integrative way to advancing our knowledge of psychiatric disorders and mental health.

Big data

Big data analytics for CLEC5A dynamics based on single cell genomics and proteomics reveal its diverse functions in human diseases.

BACKGROUND: CLEC5A (C-type lectin domain family 5 member A) is an innate immune receptor implicated in inflammatory signaling, contributing to hyperinflammatory responses in infections and sterile inflammation. However, CLEC5A dynamics in human diseases remain to be identified. Here, we systematically characterized CLEC5A dynamics in humans across cells, tissues, and disease states, and to explore the functional significance of CLEC5A in macrophage activation based on single-cell genomics. METHODS: With multi-omics (scRNA-seq, proteomics and big data analytics), we analyzed extensive human transcriptomic datasets (>42,000 samples) to profile CLEC5A expression by cell type, tissue, and disease. Single-nucleus RNA-seq (snRNA-seq) from pediatric congenital heart disease and a virtual CLEC5A gene knockout were also performed to characterize CLEC5A dynamics in humans. RESULTS: CLEC5A is highly enriched in innate immune cells, particularly in macrophages and neutrophils. Baseline CLEC5A in most tissues is low, but it is markedly upregulated in inflammatory and infectious diseases. CLEC5A expression has sex-specific differences in certain organs. Single-cell analysis showed that CLEC5A can be considered novel marker of proinflammatory macrophages with elevated cytokine production, antigen presentation, and impaired phagocytosis. Virtual CLEC5A knockout analysis identified coordinated perturbation of immune-regulatory pathways and overlapping genes linking CLEC5A to macrophage activation networks. CONCLUSION: CLEC5A is predominantly expressed in myeloid cells and acts as a key amplifier of inflammation in human diseases. Our findings highlight CLEC5A as a potential biomarker and therapeutic target in myeloid-driven hyperinflammatory conditions, warranting further experimental and translational validation.

Humans

Phylogenomic subsampling and upsampling for efficient evolutionary analyses of big data.

Long runtimes, high memory demands, and reliance on high-performance computing impede phylogenomic analyses. We review a scalable phylogenomic subsampling with upsampling (PSU) framework to address this challenge, which reduces runtime and memory requirements by orders of magnitude. In PSU, small subsamples of sites from a concatenated alignment are analyzed, which are expanded by upsampling before inference, and the resulting inferences are aggregated to obtain evolutionary estimates. PSU harnesses the fact that the computational cost of maximum likelihood analysis is strongly influenced by the number of distinct site patterns in the concatenated alignment, whereas statistical power depends primarily on the amount of evolutionary information represented by the total number of sites and substitutions. By reducing the former while restoring the latter through upsampling, PSU can approximate many full-alignment analyses at substantially lower computational cost. Analysis of simulated and empirical datasets shows that PSU can accurately estimate bootstrap support values, select the optimal substitution model, test evolutionary hypotheses, and infer branch lengths, divergence times, and associated uncertainty measures. PSU also provides distributions of inferred clade support across independent subsamples, enabling detection of conflicting phylogenetic signals that may remain hidden in conventional bootstrap analysis of concatenated alignments. Automated tuning of subsample size, the number of subsamples, and the number of upsampling replicates make PSU practical. We suggest that PSU is a general approach for scalable phylogenomic inference using a broad range of statistical methods. By enabling analyses of genome-scale alignments on commodity hardware, PSU broadens research access and reduces environmental and infrastructural costs of big-data phylogenomics.

Phylogeny

Phylogenomic subsampling and upsampling for efficient evolutionary analyses of big data.

Long runtimes, high memory demands, and reliance on high-performance computing impede phylogenomic analyses. We review a scalable phylogenomic subsampling with upsampling (PSU) framework, in which small subsamples of sites from a concatenated alignment are expanded by upsampling before inference, and the resulting analyses are then aggregated to obtain evolutionary estimates. PSU harnesses the fact that the computational cost of maximum likelihood analysis is strongly influenced by the number of distinct site patterns in the concatenated alignment, whereas statistical power depends primarily on the amount of evolutionary information represented by the total number of sites and substitutions. By reducing the former while restoring the latter through upsampling, PSU can approximate many full-data analyses at substantially lower computational cost. Analysis of simulated and empirical datasets shows that PSU can accurately estimate bootstrap support values, select the optimal substitution model, test evolutionary hypotheses, and infer branch lengths, divergence times, and associated uncertainty measures, while reducing runtime and memory requirements by orders of magnitude. PSU also provides distributions of inferred clade support across independent subsamples, enabling detection of conflicting phylogenetic signals that may remain hidden in conventional bootstrap analysis. Automated tuning of subsample size, the number of subsamples, and the number of upsampling replicates make PSU practical across diverse datasets. We suggest that PSU is a general strategy for scalable phylogenomic inference using a broad range of statistical methods. By enabling analyses of genome-scale alignments on commodity hardware, PSU broadens research access and reduces environmental and infrastructural costs of big-data phylogenomics.

confidence limits

The Computational Revolution in Natural Product Research: A Data-Driven Roadmap for Next-Generation Drug Development.

Natural products (NPs) have historically provided the foundational scaffolds for drug development, yet traditional bioprospecting faces critical limitations: high rediscovery rates, laborious isolation workflows, and substantial attrition during clinical translation. The emergence of big data technologies is fundamentally transforming this landscape, enabling a shift from serendipity-based discovery toward systematic, data-driven approaches. This review examines how the integration of artificial intelligence (AI), machine learning (ML), and multi-omics datasets is accelerating natural product research across three key domains: (1) genome mining for biosynthetic gene cluster identification using platforms such as antiSMASH, (2) cheminformatics-driven prediction of structure-activity relationships and ADMET properties, and (3) metabolomics-guided dereplication to prioritize novel bioactive scaffolds. We evaluate the convergence of genomics, metabolomics, and computational chemistry in enabling in silico lead optimization and the discovery of cryptic metabolites from previously inaccessible microbial taxa. While challenges in data standardization and scalability persist, the synergy between big data and NP research is accelerating clinical translation. Despite persistent challenges in data standardization, scalability, and equitable benefit-sharing, the convergence of big data and NP research is poised to redefine drug development. These advances position computational NP research as a cornerstone of next-generation drug development.

big data analytics

Generative AI Models in Time-Varying Biomedical Data: Scoping Review.

BACKGROUND: Trajectory modeling is a long-standing challenge in the application of computational methods to health care. In the age of big data, traditional statistical and machine learning methods do not achieve satisfactory results as they often fail to capture the complex underlying distributions of multimodal health data and long-term dependencies throughout medical histories. Recent advances in generative artificial intelligence (AI) have provided powerful tools to represent complex distributions and patterns with minimal underlying assumptions, with major impact in fields such as finance and environmental sciences, prompting researchers to apply these methods for disease modeling in health care. OBJECTIVE: While AI methods have proven powerful, their application in clinical practice remains limited due to their highly complex nature. The proliferation of AI algorithms also poses a significant challenge for nondevelopers to track and incorporate these advances into clinical research and application. In this paper, we introduce basic concepts in generative AI and discuss current algorithms and how they can be applied to health care for practitioners with little background in computer science. METHODS: We surveyed peer-reviewed papers on generative AI models with specific applications to time-series health data. Our search included single- and multimodal generative AI models that operated over structured and unstructured data, physiological waveforms, medical imaging, and multi-omics data. We introduce current generative AI methods, review their applications, and discuss their limitations and future directions in each data modality. RESULTS: We followed the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines and reviewed 155 articles on generative AI applications to time-series health care data across modalities. Furthermore, we offer a systematic framework for clinicians to easily identify suitable AI methods for their data and task at hand. CONCLUSIONS: We reviewed and critiqued existing applications of generative AI to time-series health data with the aim of bridging the gap between computational methods and clinical application. We also identified the shortcomings of existing approaches and highlighted recent advances in generative AI that represent promising directions for health care modeling.

Artificial Intelligence

Advances in tumor subclone formation and mechanisms of growth and invasion.

Tumor subclones refer to distinct cell populations within the same tumor that possess different genetic characteristics. They play a crucial role in understanding tumor heterogeneity, evolution, and therapeutic resistance. The formation of tumor subclones is driven by several key mechanisms, including the inherent genetic instability of tumor cells, which facilitates the accumulation of novel mutations; selective pressures from the tumor microenvironment and therapeutic interventions, which promote the expansion of certain subclones; and epigenetic modifications, such as DNA methylation and histone modifications, which alter gene expression patterns. Major methodologies for studying tumor subclones include single-cell sequencing, liquid biopsy, and spatial transcriptomics, which provide insights into clonal architecture and dynamic evolution. Beyond their direct involvement in tumor growth and invasion, subclones significantly contribute to tumor heterogeneity, immune evasion, and treatment resistance. Thus, an in-depth investigation of tumor subclones not only aids in guiding personalized precision therapy, overcoming drug resistance, and identifying novel therapeutic targets, but also enhances our ability to predict recurrence and metastasis risks while elucidating the mechanisms underlying tumor heterogeneity. The integration of artificial intelligence, big data analytics, and multi-omics technologies is expected to further advance research in tumor subclones, paving the way for novel strategies in cancer diagnosis and treatment. This review aims to provide a comprehensive overview of tumor subclone formation mechanisms, evolutionary models, analytical methods, and clinical implications, offering insights into precision oncology and future translational research.

Humans

Insecticide and nutrient transport in water, related to agricultural land use of a stream basin in Ontario, Canada.

Transport by stream water of insecticides and nutrients in Big Creek, Norfolk County, Ontario, Canada, was examined by combining concentrations of substrates with flow data. Big Creek has its headwaters in dairy cattle country, its central basin area is mainly devoted to tobacco growing, and its lower reaches contain mixed farming, corn and vegetables, etc., before it flows into Lake Erie. Three sampling sites were chosen to represent these 3 different land uses. Generally concentrations of substrates in the water were quite uniform at the 3 sites resulting in transport of quantities in proportion to stream flow. Certain anomalies occurred and are discussed. The 2 chief insecticides found were DDT and dieldrin. Analyses for potassium, calcium and magnesium indicated that these nutrient losses into the stream were area-rather than usage-dependent. Midseason variations in loss of nitrogen and phosphorus may be the result of agricultural practices in the three areas represented in this study. The largest quantities of all nutrients lost occurred early in the season before crops were established.

Agriculture

Exploring novel MYH7 gene variants using in silico analyses in Korean patients with cardiomyopathy.

BACKGROUND: Pathogenic variants of MYH7, which encodes the beta-myosin heavy chain protein, are major causes of dilated and hypertrophic cardiomyopathy. METHODS: In this study, we used whole-genome sequencing data to identify MYH7 variants in 397 patients with various cardiomyopathy subtypes who were participating in the National Project of Bio Big Data pilot study in Korea. We also performed in silico analyses to predict the pathogenicity of the novel variants, comparing them to known pathogenic missense variants. RESULTS: We identified 27 MYH7 variants in 41 unrelated patients with cardiomyopathy, consisting of 20 previously known pathogenic/likely pathogenic variants, 2 variants of uncertain significance, and 5 novel variants. Notably, the pathogenic variants predominantly clustered within the myosin motor domain of MYH7. We confirmed that the novel identified variants could be pathogenic, as indicated by high prediction scores in the in silico analyses, including SIFT, Mutation Assessor, PROVEAN, PolyPhen-2, CADD, REVEL, MetaLR, MetaRNN, and MetaSVM. Furthermore, we assessed their damaging effects on protein dynamics and stability using DynaMut2 and Missense3D tools. CONCLUSIONS: Overall, our study identified the distribution of MYH7 variants among patients with cardiomyopathy in Korea, offering new insights for improved diagnosis by enriching the data on the pathogenicity of novel variants using in silico tools and evaluating the function and structural stability of the MYH7 protein.

Humans

Microbial partnerships and molecular mechanisms in plant stress physiology for climate-resilient and sustainable farming.

Plant-microbial partnerships and their underlying molecular mechanisms are indispensable, natural drivers of improved nutrient acquisition and stress tolerance in the face of climate-driven environmental challenges. Modern multi-omics tools, when coupled with artificial intelligence and synthetic biology, enable the precise design of targeted bioinoculants and synthetic microbial consortia. Translating these advanced microbiome-based strategies into scalable, field-level agricultural applications provides a sustainable path toward securing global food production while maintaining soil health. Global climate change imposes multifaceted abiotic and biotic stresses on crops, disrupting physiological and molecular processes and threatening agricultural productivity. Plant-associated microbes represent an underexplored yet powerful ally in enhancing crop resilience. This review presents current knowledge of plant-microbe interactions and the molecular mechanisms governing plant stress physiology, with an emphasis on climate-resilient and sustainable farming. Hence, ever-changing environmental cues pose a significant burden on agricultural productivity, and plant-associated microbial communities modulate a cascade of physiological and molecular responses, including production of phytohormones, signaling, regulation of reactive oxygen species homeostasis, and activation of plant immune responses to help plants withstand stress and enhance productivity. Moreover, root exudates, phytohormones, and quorum sensing mediate the central communication networks, facilitating plant-microbe cross talk. Additionally, the advances in OMICs approaches aid in disentangling the molecular underpinnings of these interactions by providing mechanistic insights and potential candidate gene targets for crop improvement and stress resilience. In the post-genomic era, integrating artificial intelligence and big data analysis to optimize microbiome-based strategies for sustainable agriculture is a new frontier for disentangling plant-microbe symbiosis to improve soil health, enhance crop yields, and improve stress tolerance. Thus, by integrating the ecological, physiological, and molecular perspectives, this review highlights the transformative potential of harnessing plant-microbe symbiosis for climate-resilient and sustainable agriculture.

Stress, Physiological

Dissemination of antimicrobial resistance in Klebsiella spp. from urban aquatic environments: a multi-country genomic perspective.

INTRODUCTION: Antibiotic resistance, particularly carbapenem-resistant Klebsiella pneumoniae (CRKP), poses significant clinical and environmental threats, especially in urban aquatic ecosystems and hospital wastewaters. OBJECTIVES: This study aims to analyze the epidemiological and genomic features of CRKP isolates in urban aquatic environments and evaluate their public health and environmental impacts. METHODS AND RESULTS: Water samples were collected from 113 rivers and 3 hospitals in China, Sri Lanka, and Nepal to isolate carbapenem-resistant Klebsiella spp. isolates. Antimicrobial susceptibility testing, whole-genome sequencing, and bioinformatics analyses were performed to characterize resistance phenotypes, antibiotic resistance genes (ARGs), and evolutionary trends. Big data analysis further elucidated the genomic characteristics of CRKP in global water sources, and Galleria mellonella larvae were used to assess virulence. Statistical analysis validated the findings. A total of 192 carbapenem-resistant Klebsiella spp. isolates were identified from urban aquatic ecosystems in China (n = 60) and Nepal (n = 132), with CRKP (n = 161) being the predominant species. All CRKP isolates exhibited a multidrug-resistant phenotype, yet significant differences in resistance profiles and associated ARGs were observed between isolates from the two countries. Nine carbapenem resistance genes (CRGs) were detected, with blaNDM-1 being the most prevalent (57.8 %). Correlation analysis revealed a strong association between these CRGs and multiple Inc-type plasmids. Global genomic analysis of CRKP from water sources across eight countries identified ten distinct CRGs across 45 serotypes, with KL64 being the most predominant. Notably, carbapenem-resistant hypervirulent Klebsiella pneumoniae was detected in water samples from Nepal. CONCLUSION: Our findings highlight significant regional disparities in CRKP prevalence and ARG dissemination across urban aquatic environments, with Nepal showing the highest prevalence, particularly in untreated rivers. China exhibited lower prevalence but distinct resistance gene profiles, while no CRKP was detected in Sri Lanka, underscoring the impact of environmental management and healthcare infrastructure on ARG spread.

Humans

Proteomics at scale: Bottlenecks and opportunities for early-career researchers in a fast developing field.

The field of proteomics has rapidly evolved over the last five years enabled by rapid advances in instrumentation and computation. At the same time, the proteomics community is also growing. This is reflected by the increasing participation in international conferences such as those organized by the European Proteomics Association and the Human Proteome Organization. These events provide early-career researchers with unique opportunities to exchange ideas, develop collaborations, and build networks that support professional development. One such network is the Young Proteomics Investigators Club, a European initiative supported by European Proteomics Association and led by early-career researchers. In this Community-Driven project, we investigate recent trends in proteomics by screening conference abstracts and evaluating the session attendance at Human Proteome Organization Congresses and European Proteomics Association conferences. Based on these analyses, we identified five areas that, from our perspective, are shaping the current trends in proteomics: clinical proteomics, proteomics of post-translational modifications, single-cell proteomics, systems biology and multi-omics, and computational proteomics. For each area, we highlight both unique challenges and identify a common theme: a shift from exploratory studies with manageable sample numbers towards large screenings and cohorts and the generation of big data, which often comes with the lack of computational support, organizational networks, and infrastructure. In this light, we describe the unique challenges and opportunities faced by early-career researchers. We point to actionable directions for enabling reproducible and transparent proteomics as well as community-driven projects and initiatives, which are often providing training and support. SIGNIFICANCE: In this perspective, the Young Proteomics Investigators Club (YPIC) discusses advances in analytical developments and computational approaches in proteomics research. Based on empirical analysis of recent European Proteomics Association conference and Human Proteome Organization congresses contributions, we identify clinical, single-cell, post-translational and systems-level proteomics as the research areas that have gained most momentum in the last three to five years. What makes this work distinctive is that it is written by and for early-career researchers, thereby uniquely identifying where momentum, challenges, and unmet needs converge for the newest generation of proteomics researchers. Rather than cataloguing advances, we examine the widening gap between what modern proteomics can generate and what individual researchers can realistically process, validate, and interpret. We describe specific structural barriers including access to high performance computing, limited formal training in scalable data analysis, the need for unified benchmarking standards and navigating clinical collaboration frameworks. We then highlight opportunities for the field, such as community-curated benchmarks, interdisciplinary mentorship models, and shared computational infrastructure. By making these challenges explicit from an early-career researchers standpoint, we aim to inform how training, funding, and community initiatives can be shaped to support the next generation of proteomics researchers.

Proteomics

Resilience indicator traits in chickens: a systematic review.

The resilience of an animal is its ability to cope with short-term disturbances, including those caused by pathogens, through response and rapid recovery to its original state. Here, a systematic literature review was conducted to identify resilience indicators that have already been studied and implemented for chickens, along with the contexts or specific stressor under which they were examined. The literature review was based on predefined search criteria, including 'chickens' as the population of interest, a stress exposure description and the term 'resilience'. Screening of titles and abstracts was assisted by the AI-based tool ASReview, followed by full text screening and analysis. According to selection criteria, we finally identified 33 relevant publications on resilience indicator traits used in chickens. Two of these studies analyzed chicken resilience based on routinely collected big data, while the remaining majority were small or medium scale studies and trials. The majority of the studies focused on immune (n = 13) or thermal challenges (either before or after hatching; n = 14), especially heat stress. A variety of resilience indicators were studied; most studies investigated production- or performance-related parameters and/or immunity or disease-related parameters. About half of the studies included genetics- or gene expression-related indicators. Investigation of behavioral indicators of resilience - other than feed intake - was limited (n = 4). Overall, this systematic review provides a comprehensive overview of published resilience indicators in chickens and highlights gaps for future research. This review also revealed a need for the controlled use of terms like resilience, robustness, resistance, tolerance or adaptability, and the relevance of assessing the phase of recovery after a short-term disturbance in the context of resilience.

Animals

Discovery of novel diagnostic biomarkers of hepatocellular carcinoma associated with immune infiltration.

OBJECTIVE: Diagnosis of hepatocellular carcinoma (HCC) remains challenging for clinicians. Machine learning approaches and big data analyses are viable strategies for identifying HCC diagnostic markers. MATERIALS AND METHODS: In this study, we downloaded mRNA expression profiles of HCC from the GEO database and used random forest and machine learning algorithms, such as least absolute shrinkage and selection operator, to screen for reliable diagnostic genes. Disease Ontology, Kyoto Encyclopedia of Genes and Genomes (KEGG) and Gene Set Enrichment Analysis enrichment analyses were performed to explore differential gene functions and disease pathways. CIBERSORT was performed to calculate the immune cell infiltration of HCC and the correlation between diagnostic genes and immune cells. Cell experiments were performed to evaluate the function of R-spondin 3 (RSPO3) in HCC cells. Immunohistochemical staining was used to evaluate the protein expression of CD138, CD206 and iNOS. RESULTS: The results indicated that extracellular matrix protein 1 (ECM1), Niemann-Pick C1-Like 1 (NPC1L1) and RSPO3 were down-regulated in HCC compared with the normal group (p&#x2009;<&#x2009;0.05), which was validated in clinical tissue samples. Moreover, ECM1, NPC1L1 and RSPO3 had high diagnostic values (AUC > 0.75) for HCC in both training and test groups. Immuno-infiltration analysis revealed that ECM1 and RSPO3 were highly positively correlated with neutrophil and macrophage M2 levels, whereas they were negatively correlated with Tregs. RSPO3-si affected cell proliferation and apoptosis in HCC. Furthermore, RSPO3 exhibited a positive correlation with tumour progression, the proportion of plasma cells and M2 macrophages in mice, while showing a negative association with M1 macrophages. CONCLUSION: The present study identified ECM1, NPC1L1 and RSPO3 as new diagnostic biomarkers for HCC based on normal and diseased samples from HCC, meanwhile the pro-oncogenic function of RSPO3 and its regulation on immune infiltration have been confirmed.

Carcinoma, Hepatocellular

Interpreting cancer genetics through a two-step "evolutionary cascade hypothesis": bridging neutral and selective perspectives.

BACKGROUND: DNA mutations are the fundamental engines of cancer, driving its initiation and progression. The forces that fuel malignancy are also the architects of evolution, shaping life through genetic variations. Mutations, in fact, can emerge naturally from endogenous processes, such as oxidative DNA damage or errors in replication, as well as induced by external factors, including cosmic radiation and chemical carcinogens. MAIN BODY: A key question in cancer research is whether tumor evolution is primarily governed by selective bottlenecks, neutral evolution, or dynamic genetic plasticity. In this work, we examine cancer as a disease driven by evolutionary processes rooted in fundamental biological requirements, including sustained proliferation and nutrient utilization. We hypothesize that the accumulation of mutations activates an evolutionary switch, enabling tumor cells to acquire an enhanced capacity for survival, adaptation, and growth at rates far exceeding typical evolutionary timescales. We propose the "evolutionary cascade hypothesis," a unifying framework that integrates these models into a coherent sequence. At its core lies the failure of DNA repair mechanisms, representing a critical transition in cancer progression. This shift marks the transition from an initial non-Darwinian, neutral phase to a Darwinian, more deterministic phase. CONCLUSIONS: As predictive models of tumor evolution advance through genomic big data and artificial intelligence-driven analysis, the future of cancer treatment may extend beyond targeting individual mutations to disrupting the underlying evolutionary mechanisms that sustain malignancy. This paradigm shift could redefine therapeutic strategies and ultimately improve patient outcomes.

Humans

Streamlining large-scale genomic data management: Insights from the UK Biobank whole-genome sequencing data.

Biobank-scale whole-genome sequencing (WGS) studies are increasingly pivotal in unraveling the genetic bases of diverse health outcomes. However, managing and analyzing these datasets' sheer volume and complexity presents significant challenges. We highlight the annotated genomic data structure (aGDS) format, substantially reducing the WGS data file size while enabling seamless integration of genomic and functional information for comprehensive WGS analyses. The aGDS format yielded 23 chromosome-specific files for the UK Biobank 500k WGS dataset, occupying only 1.10 tebibytes of storage. We develop the vcf2agds toolkit that streamlines the conversion of WGS data from VCF to aGDS format. Additionally, the STAARpipeline equipped with the aGDS files enabled scalable, comprehensive, and functionally informed WGS analysis, facilitating the detection of common and rare coding and noncoding phenotype-genotype associations. Overall, the vcf2agds toolkit and STAARpipeline provide a streamlined solution that facilitates efficient data management and analysis of biobank-scale WGS data across hundreds of thousands of samples.

Humans