Search PubMedSearch

SEARCH · Search PubMed

Results for “Big Data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

29 records · Page 2Linked to original sources

Machine learning for population-level risk prediction of future cholangiocarcinoma.

BACKGROUND: The poor prognosis of cholangiocarcinoma (CCA) is largely driven by rapid, asymptomatic disease progression, which usually results in a late diagnosis in the absence of established screening strategies. An early, cost-effective, and universally applicable risk assessment strategy would therefore be valuable. METHODS: We developed machine learning (ML) models on prospective, multimodal data from 487,495 UK Biobank (UKB) participants, of whom 649 developed CCA during follow-up. Data from England (80%) were utilised for ML development via five-fold cross-validation, and then all models were tested on withheld data from Scotland, Wales, and Newcastle (20%). Iterative ablation studies reduced inputs from >150 features across demographic data, lifestyle, health records, blood parameters, genomics, and metabolomics to models built on five and ten routinely available clinical parameters. These were externally validated in the Penn Medicine Biobank (PMBB; n = 2638; 28 CCA), All of Us Research Program (AOU; n = 330,433; 362 CCA), Japan Medical Data Centre Claims Database (JMDC; n = 8,425,522; 723 CCA) and TriNetX (n = 728,886; 1592 CCA). FINDINGS: We show that ML models integrating biliary-disease associated health records and Gamma glutamyltransferase can stratify risk of future CCA. Evaluation on the UKB test set as well as three independent cohorts revealed robust performance and generalisability across ethnicities. We achieved AUROCs of 0.71 [95% CI: 0.703-0.711], 0.77 [95% CI: 0.764-0.778 ], 0.796 [95% CI: 0.795-0.798] and 0.8 [95% CI: 0.794-0.805] for UKB, PMBB, AOU, and JMDC respectively, with respective AUPRCs of 0.014 [95% CI: 0.009-0.018], 0.042 [95% CI: 0.037-0.048], 0.038 [95% CI: 0.033-0.042] and 0.001 [95% CI: 0.001-0.001]. In AOU, application of the Youden J-optimised threshold yielded a number needed to screen of 79. Separate models for intra- and extrahepatic CCA did not improve performance. In line with the pathophysiology, performance declined for longer intervals between assessment and event. A group-level analysis in the TriNetX cohort revealed hazard ratios of up to 82.5 [95% CI: 26.4-257.96]. We provide extensive interpretability results and release all source codes used to develop the presented models. INTERPRETATION: We provide a comprehensive framework for early CCA risk stratification in the general population, identifying key predictors, and demonstrating the potential of data-driven models in personalised screening for hepatobiliary cancer. FUNDING: German Cancer Aid (grant #70115730), Junior Principal Investigator Fellowship programme of RWTH Aachen Excellence strategy.

Humans

Evolution and applications of genome-scale metabolic models in yeast systems biology studies.

Genome-scale metabolic models (GEMs) can be used to simulate the metabolic network of an organism in a systematic and holistic way. Different yeast species, including Saccharomyces cerevisiae, have emerged as powerful cell factories for bioproduction. Recently, with the dedicated efforts from the scientific community, significant progress has been made in the development of yeast GEMs. Numerous versions of yeast GEMs and the derived multiscale models have been released, facilitating integrative omics analysis and rational strain design for different types of yeast cell factories. These advancements reflected the evolution and maturation of yeast GEMs together with a model ecosystem around them. This review will summarize the development and expansion of yeast GEMs and discuss their applications in yeast systems biology studies. It is anticipated that yeast GEMs will continue to play an increasingly important role in pioneering yeast physiological and metabolic studies in coming years.

Systems Biology

Through the lens of bioenergy crops: advances, bottlenecks, and promises of plant engineering.

Advances in engineering of bioenergy crops were driven over the past years by adapting technological breakthroughs and accelerating conventional applications but also exposed intriguing challenges. New tools revealed rich interconnectivity in the exponentially growing and dynamic 'big' omics data' of metabolomes, transcriptomes, and genomes at previously inaccessible magnitude (global, cross-species, meta-) and resolution (single cell). Insights enabled fresh hypotheses and stimulated disciplines such as functional genomics with discovery of broad regulatory networks and their determinants, that is, DNA parts, including promoters, regulatory elements, and transcription factors. Their rational design, assembly into increasingly complex blueprints, and installation into diverse chassis is an existing frontier that may benefit from emerging technologies to address bottlenecks. Interweaving nature-inspired to fully synthetic parts has already allowed building of fine-tuned regulatory circuits, or new-to-nature metabolic routes insulated from the biological context of the chassis species. Similarly, developments and the evolving need for unifying principles in plant transformation and species-agnostic technologies highlight future opportunities for engineering the next generation of bioenergy plants.

Crops, Agricultural

Robust inference and correlates from genetic associations with personality.

Personality traits describe stable differences in how people think, feel and behave, and how they interact with and experience their social and physical environments1,2. Many questions remain unanswered about associations between DNA and personality traits, such as their robustness, their generalizability and the biological and social pathways through which they act. Here we meta-analyse data across 46 cohorts comprising 611,037 to 1.14 million participants with European-like and African-like genomes for genome-wide association studies (GWAS) of the Big Five personality traits (extraversion, agreeableness, conscientiousness, neuroticism and openness to experience), and data from up to 50,725 participants for within-family GWAS. We identify 1,260 lead genetic variants associated with personality, including 824 novel variants3. Common genetic variants explain a moderate 4.8-9.3% of the variance in measures of each trait, and 9.3-13.3% among instruments with typical measurement reliability. Genetic associations with personality are highly consistent but not identical across geography, reporter (self versus close other), age group and measurement instrument, and we find minimal spousal assortment for personality in recent history. In contrast to many other social and behavioural traits4,5, within-family GWAS and polygenic index analyses indicate that genetic associations with personality are minimally confounded by the shared family environment. Polygenic prediction, genetic correlation and Mendelian randomization analyses indicate that personality traits have widespread, potentially causal associations with consequential behaviours and life outcomes. Overall, we find that the genetic architecture of personality is robustly generalizable, minimally confounded and widely relevant to human experience.

Journal Article

Benchmarking methods for measuring biosynthetic gene cluster similarity and determination of gene cluster families.

MOTIVATION: Natural products are often produced by a set of biosynthetic enzymes that are encoded by genes clustered together in the producer's genome, referred to as a biosynthetic gene cluster (BGC). The ability to compare and cluster BGCs is essential for several applications, including predicting which bacteria will make a known product and assessing the potential diversity of natural products produced by a set of bacteria. There are multiple methods for comparing and clustering BGCs based on their similarity, but there has been a lack of investigation into how strongly BGC similarity relates to product structural similarity and how these methods perform relative to each other. RESULTS: Using publicly available databases, we developed a benchmark dataset to assess how well different BGC similarity metrics correlate with the structural similarity of their products and how well these methods cluster BGCs. We found that all methods showed moderate correlation between BGC and structural similarity, with correlations improving for more similar BGCs and varying significantly by BGC biosynthetic class. Analysis of outliers revealed some outliers were due to mistakes or omissions in public datasets, while others represented deviation between BGC similarity and product structural similarity. All methods generally performed better on clustering metrics, with BiG-SCAPE performing the best after errors in the public datasets had been corrected. AVAILABILITY AND IMPLEMENTATION: Scripts and data required to reproduce the results are available at https://github.com/aswalker-lab/BGC-clustering-benchmark and processed similarity, clusters, and scaffolds are also available at https://huggingface.co/datasets/allie-walker/BGC-clustering-benchmark. Code is also available at Zenodo: 10.5281/zenodo.17373546.

Multigene Family

A large-scale study across the avian clade identifies ecological drivers of neophobia.

Neophobia, or aversion to novelty, is important for adaptability and survival as it influences the ways in which animals navigate risk and interact with their environments. Across individuals, species and other taxonomic levels, neophobia is known to vary considerably, but our understanding of the wider ecological drivers of neophobia is hampered by a lack of comparative multispecies studies using standardized methods. Here, we utilized the ManyBirds Project, a Big Team Science large-scale collaborative open science framework, to pool efforts and resources of 129 collaborators at 77 institutions from 24 countries worldwide across six continents. We examined both difference scores (between novel object test and control conditions) and raw data of latency to touch familiar food in the presence (test) and absence (control) of a novel object among 1,439 subjects from 136 bird species across 25 taxonomic orders incorporating lab, field, and zoo sites. We first demonstrated that consistent differences in neophobia existed among individuals, among species, and among other taxonomic levels in our dataset, rejecting the null hypothesis that neophobia is highly plastic at all taxonomic levels with no evidence for evolutionary divergence. We then tested for effects of ecological factors on neophobia, including diet, sociality, habitat, and range, while accounting for phylogeny. We found that (i) species with more specialist diets were more neophobic than those with more generalist diets, providing support for the Neophobia Threshold Hypothesis; (ii) migratory species were also more neophobic than nonmigratory species, which supports the Dangerous Niche Hypothesis. Our study shows that the evolution of avian neophobia has been shaped by ecological drivers and demonstrates the potential of Big Team Science to advance our understanding of animal behavior.

Animals

Expression regulation network in papillae of sea cucumbers: Whole-transcriptome and DNA methylation datasets.

To elucidate the expression regulation network of papilla size of sea cucumbers (Apostichopus japonicus), the whole-transcriptome and DNA methylome datasets of different sizes of papillae in sea cucumbers were generated. Average clean bases of whole-transcriptome (16.35 G) and DNA methylome (28.92 G) were obtained using RNA sequencing and whole-genome bisulfite sequencing techniques. A total of 3,188 ceRNA networks were also identified including 3,081 long non-coding RNAs (lncRNA)/microRNAs (miRNA)/mRNA networks and 107 circular RNA (circRNA)/miRNA/mRNA networks. Methylome data indicate that there were 3,307 and 3,776 differentially methylated regions (DMRs) with high-level methylation as well as 3,125 and 3,016 DMRs with low-level methylation in big papillae compared to small papillae. The identified DMRs were mainly distributed in introns, promotors, or exons. The whole-transcriptome and DNA methylome datasets generated from this study not only established a robust theoretical foundation (especially from the epigenetic aspect) for elucidating expression regulation network determining papilla size in sea cucumbers but also can be a valuable resource of biomarker mining for papilla appearance-based selective breeding in sea cucumbers.

DNA Methylation

Phoronida-A small clade with a big role in understanding the evolution of lophophorates.

Phoronids, together with brachiopods and bryozoans, form the animal clade Lophophorata. Modern lophophorates are quite diverse-some can biomineralize while others are soft-bodied, they could be either solitary or colonial, and they develop through various eccentric larval stages that undergo different types of metamorphoses. The diversity of this clade is further enriched by numerous extinct fossil lineages with their own distinct body plans and life histories. In this review, I discuss how data on phoronid development, genetics, and morphology can inform our understanding of lophophorate evolution. The actinotrocha larvae of phoronids is a well documented example of intercalation of the new larval body plan, which can be used to study how new life stages emerge in animals with biphasic life cycle. The genomic and embryonic data from phoronids, in concert with studies of the fossil lophophorates, allow the more precise reconstruction of the evolution of lophophorate biomineralization. Finally, the regenerative and asexual abilities of phoronids can shed new light on the evolution of coloniality in lophophorates. As evident from those examples, Phoronida occupies a central role in the discussion of the evolution of lophophorate body plans and life histories.

Animals

Sex-biased Migration and Demographic History of the Big European Firefly Lampyris noctiluca.

Differential dispersion between the sexes can impact the colonization process and demographic history of a species. Here, we explored the demographic history of the big European firefly, Lampyris noctiluca, which exhibits female neoteny. Distribution of L. noctiluca extends throughout Europe, but nothing is known about its colonization process. To investigate its demographic history, we produced the first Lampyris genome (653 Mb), including an IsoSeq annotation and the identification of the X chromosome. We collected 115 individuals from six populations of L. noctiluca (Finland to Italy) and generated whole-genome re-sequencing data for each individual. We inferred several population expansions and bottlenecks throughout the Pleistocene that correlate with glaciation events. Surprisingly, we uncovered strong population structure and low gene flow. We reject a stepwise, south to north, colonization history scenario and instead uncovered a complex demographic history with a putative eastern European origin. Analyzing the evolutionary history of the mitochondrial genome as well as X-linked and autosomal loci, we found evidence of a maternal colonialization of Germany, putatively from a farther western European population, followed by a male-only migration from south of the Alps (Italy). Overall, investigating the demographic history and colonization patterns of a species should form part of an integrative approach of biodiversity research. Our results provide evidence of sex-biased migration which is important to consider for demographic, biogeographic and species delimitation studies.

Animals

Differential Mutagenic Response of Rat Liver and Lung to Nicotine-Derived Nitrosamine Ketone (NNK).

Nitrosamines (NA) are chemical impurities that are present in tobacco, foods, more recently in some pharmaceuticals and are associated with genotoxicity and carcinogenicity. We evaluated the in vivo mutagenicity of nicotine-derived nitrosamine ketone (NNK) or 4-(methyl nitrosamino)-1-(3-pyridyl)-1-butanone, a model compound used as an anchor molecule to estimate carcinogenic potency of unknown nitrosamine impurities. Big Blue rats were treated with NNK at doses ranging from 0.001 to 30 mg/kg for 28 days, following which liver and lung tissue were harvested 3 days later for nuclear genomic DNA isolation. Mutations in liver and lung were assessed with the cII transgene assay and endogenous genomic loci using Duplex Sequencing (DupSeq), a highly validated error-corrected sequencing (ECS) technology. The no genotoxic effect level (NOGEL) was 1 mg/kg in liver and 0.1 mg/kg in lung while the benchmark dose (BMD) analysis for cII mutagenicity determined a BMDL50 of 1.3 mg/kg in liver and 0.12 mg/kg in lung, consistent with lung being the more sensitive target organ for carcinogenicity for NNK. ECS-derived mutagenicity was highly correlated with cII-derived mutagenicity. Interestingly, the types of mutations formed appeared to be tissue-specific with higher C > T transitions and lower T > G transversions in lung compared to liver, differences that may reflect tissue-specific DNA repair capacity and/or metabolic differences. Collectively, these data support the use of in vivo mutagenicity data─from both TGR cII and ECS methods─for human health and cancer risk characterization of nitrosamines and for estimating acceptable daily intakes for unknown nitrosamine drug substance related impurities.

Animals

Leading with Innovation: Maternal Health Transformation in New York City Health + Hospitals.

New York City's (NYC) maternal health crisis drew close attention in the late 2010s, driven by alarming data: Approximately 30 women died annually during childbirth in NYC, Black non-Hispanic women were 12 times more likely to die than white women, and more than 3,000 women experienced life-threatening birth complications each year. In response, NYC committed $12.8 million in July 2018 to reduce maternal mortality and eliminate racial disparities.NYC Health + Hospitals (H+H)-the nation's largest public health system, serving 1.1 million patients annually with roughly 15,000 births per year-became the primary vehicle for this initiative. With 80 percent of the system's deliveries covered by Medicaid and a patient population that is 51.2 percent Hispanic and 27.1 percent Black, H+H is uniquely positioned to lead the fight against maternal health inequity.Three flagship programs anchor H+H's response to the city's maternal mortality rate. The OB Simulation Program, launched in 2012 and expanded in 2018, was the first in the nation to use mannequins of color to train thousands of providers in obstetric emergencies. The Maternal Home Program, piloted at H+H's Kings County Hospital in 2019 and scaled system-wide by 2021, has served more than 10,341 patients, generating more than 33,000 referrals for social, behavioral health, and community resources. The Cardio-Obstetrics Program located at Kings County Hospital targets cardiovascular disease-the leading cause of maternal death among Black women-through screening, education, and community outreach. These programs are a health equity imperative, made more urgent by impending federal Medicaid cuts resulting from the H.R.1 One Big Beautiful Bill Act (passed on July 4, 2025).

Humans