Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “big data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Dissemination of antimicrobial resistance in Klebsiella spp. from urban aquatic environments: a multi-country genomic perspective.

INTRODUCTION: Antibiotic resistance, particularly carbapenem-resistant Klebsiella pneumoniae (CRKP), poses significant clinical and environmental threats, especially in urban aquatic ecosystems and hospital wastewaters. OBJECTIVES: This study aims to analyze the epidemiological and genomic features of CRKP isolates in urban aquatic environments and evaluate their public health and environmental impacts. METHODS AND RESULTS: Water samples were collected from 113 rivers and 3 hospitals in China, Sri Lanka, and Nepal to isolate carbapenem-resistant Klebsiella spp. isolates. Antimicrobial susceptibility testing, whole-genome sequencing, and bioinformatics analyses were performed to characterize resistance phenotypes, antibiotic resistance genes (ARGs), and evolutionary trends. Big data analysis further elucidated the genomic characteristics of CRKP in global water sources, and Galleria mellonella larvae were used to assess virulence. Statistical analysis validated the findings. A total of 192 carbapenem-resistant Klebsiella spp. isolates were identified from urban aquatic ecosystems in China (n = 60) and Nepal (n = 132), with CRKP (n = 161) being the predominant species. All CRKP isolates exhibited a multidrug-resistant phenotype, yet significant differences in resistance profiles and associated ARGs were observed between isolates from the two countries. Nine carbapenem resistance genes (CRGs) were detected, with blaNDM-1 being the most prevalent (57.8 %). Correlation analysis revealed a strong association between these CRGs and multiple Inc-type plasmids. Global genomic analysis of CRKP from water sources across eight countries identified ten distinct CRGs across 45 serotypes, with KL64 being the most predominant. Notably, carbapenem-resistant hypervirulent Klebsiella pneumoniae was detected in water samples from Nepal. CONCLUSION: Our findings highlight significant regional disparities in CRKP prevalence and ARG dissemination across urban aquatic environments, with Nepal showing the highest prevalence, particularly in untreated rivers. China exhibited lower prevalence but distinct resistance gene profiles, while no CRKP was detected in Sri Lanka, underscoring the impact of environmental management and healthcare infrastructure on ARG spread.

Humans↗

Proteomics at scale: Bottlenecks and opportunities for early-career researchers in a fast developing field.

The field of proteomics has rapidly evolved over the last five years enabled by rapid advances in instrumentation and computation. At the same time, the proteomics community is also growing. This is reflected by the increasing participation in international conferences such as those organized by the European Proteomics Association and the Human Proteome Organization. These events provide early-career researchers with unique opportunities to exchange ideas, develop collaborations, and build networks that support professional development. One such network is the Young Proteomics Investigators Club, a European initiative supported by European Proteomics Association and led by early-career researchers. In this Community-Driven project, we investigate recent trends in proteomics by screening conference abstracts and evaluating the session attendance at Human Proteome Organization Congresses and European Proteomics Association conferences. Based on these analyses, we identified five areas that, from our perspective, are shaping the current trends in proteomics: clinical proteomics, proteomics of post-translational modifications, single-cell proteomics, systems biology and multi-omics, and computational proteomics. For each area, we highlight both unique challenges and identify a common theme: a shift from exploratory studies with manageable sample numbers towards large screenings and cohorts and the generation of big data, which often comes with the lack of computational support, organizational networks, and infrastructure. In this light, we describe the unique challenges and opportunities faced by early-career researchers. We point to actionable directions for enabling reproducible and transparent proteomics as well as community-driven projects and initiatives, which are often providing training and support. SIGNIFICANCE: In this perspective, the Young Proteomics Investigators Club (YPIC) discusses advances in analytical developments and computational approaches in proteomics research. Based on empirical analysis of recent European Proteomics Association conference and Human Proteome Organization congresses contributions, we identify clinical, single-cell, post-translational and systems-level proteomics as the research areas that have gained most momentum in the last three to five years. What makes this work distinctive is that it is written by and for early-career researchers, thereby uniquely identifying where momentum, challenges, and unmet needs converge for the newest generation of proteomics researchers. Rather than cataloguing advances, we examine the widening gap between what modern proteomics can generate and what individual researchers can realistically process, validate, and interpret. We describe specific structural barriers including access to high performance computing, limited formal training in scalable data analysis, the need for unified benchmarking standards and navigating clinical collaboration frameworks. We then highlight opportunities for the field, such as community-curated benchmarks, interdisciplinary mentorship models, and shared computational infrastructure. By making these challenges explicit from an early-career researchers standpoint, we aim to inform how training, funding, and community initiatives can be shaped to support the next generation of proteomics researchers.

Proteomics↗

Resilience indicator traits in chickens: a systematic review.

The resilience of an animal is its ability to cope with short-term disturbances, including those caused by pathogens, through response and rapid recovery to its original state. Here, a systematic literature review was conducted to identify resilience indicators that have already been studied and implemented for chickens, along with the contexts or specific stressor under which they were examined. The literature review was based on predefined search criteria, including 'chickens' as the population of interest, a stress exposure description and the term 'resilience'. Screening of titles and abstracts was assisted by the AI-based tool ASReview, followed by full text screening and analysis. According to selection criteria, we finally identified 33 relevant publications on resilience indicator traits used in chickens. Two of these studies analyzed chicken resilience based on routinely collected big data, while the remaining majority were small or medium scale studies and trials. The majority of the studies focused on immune (n = 13) or thermal challenges (either before or after hatching; n = 14), especially heat stress. A variety of resilience indicators were studied; most studies investigated production- or performance-related parameters and/or immunity or disease-related parameters. About half of the studies included genetics- or gene expression-related indicators. Investigation of behavioral indicators of resilience - other than feed intake - was limited (n = 4). Overall, this systematic review provides a comprehensive overview of published resilience indicators in chickens and highlights gaps for future research. This review also revealed a need for the controlled use of terms like resilience, robustness, resistance, tolerance or adaptability, and the relevance of assessing the phase of recovery after a short-term disturbance in the context of resilience.

Animals↗

The mean angular distance among objects and its relationships with Kohonen artificial neural networks.

This job refers to classification of multidimensional objects and Kohonen artificial neural networks. A new concept is introduced, called the mean angular distance among objects (MADO). Its value can be calculated as the cosine of the mean centered vectors between objects. It can be expressed in matrix form for any number of objects. The MADO allows us to interpret the final organization of the objects in a Kohonen map. Simulated examples demonstrate the relationship between MADO and Kohonen maps and show a way to take advantage of the information present in both of them. Finally, a real analytical chemistry case is analyzed as an application on a big data set of an air quality monitoring campaign. It is possible to discover in it a subgroup of objects with different characteristics than those of the general trend. This subgroup is linked to the existence of an unidentified SO(2) source that, a priori, has not been taken into account.

Journal Article↗

Discovery of novel diagnostic biomarkers of hepatocellular carcinoma associated with immune infiltration.

OBJECTIVE: Diagnosis of hepatocellular carcinoma (HCC) remains challenging for clinicians. Machine learning approaches and big data analyses are viable strategies for identifying HCC diagnostic markers. MATERIALS AND METHODS: In this study, we downloaded mRNA expression profiles of HCC from the GEO database and used random forest and machine learning algorithms, such as least absolute shrinkage and selection operator, to screen for reliable diagnostic genes. Disease Ontology, Kyoto Encyclopedia of Genes and Genomes (KEGG) and Gene Set Enrichment Analysis enrichment analyses were performed to explore differential gene functions and disease pathways. CIBERSORT was performed to calculate the immune cell infiltration of HCC and the correlation between diagnostic genes and immune cells. Cell experiments were performed to evaluate the function of R-spondin 3 (RSPO3) in HCC cells. Immunohistochemical staining was used to evaluate the protein expression of CD138, CD206 and iNOS. RESULTS: The results indicated that extracellular matrix protein 1 (ECM1), Niemann-Pick C1-Like 1 (NPC1L1) and RSPO3 were down-regulated in HCC compared with the normal group (p&#x2009;<&#x2009;0.05), which was validated in clinical tissue samples. Moreover, ECM1, NPC1L1 and RSPO3 had high diagnostic values (AUC > 0.75) for HCC in both training and test groups. Immuno-infiltration analysis revealed that ECM1 and RSPO3 were highly positively correlated with neutrophil and macrophage M2 levels, whereas they were negatively correlated with Tregs. RSPO3-si affected cell proliferation and apoptosis in HCC. Furthermore, RSPO3 exhibited a positive correlation with tumour progression, the proportion of plasma cells and M2 macrophages in mice, while showing a negative association with M1 macrophages. CONCLUSION: The present study identified ECM1, NPC1L1 and RSPO3 as new diagnostic biomarkers for HCC based on normal and diseased samples from HCC, meanwhile the pro-oncogenic function of RSPO3 and its regulation on immune infiltration have been confirmed.

Carcinoma, Hepatocellular↗

Interpreting cancer genetics through a two-step "evolutionary cascade hypothesis": bridging neutral and selective perspectives.

BACKGROUND: DNA mutations are the fundamental engines of cancer, driving its initiation and progression. The forces that fuel malignancy are also the architects of evolution, shaping life through genetic variations. Mutations, in fact, can emerge naturally from endogenous processes, such as oxidative DNA damage or errors in replication, as well as induced by external factors, including cosmic radiation and chemical carcinogens. MAIN BODY: A key question in cancer research is whether tumor evolution is primarily governed by selective bottlenecks, neutral evolution, or dynamic genetic plasticity. In this work, we examine cancer as a disease driven by evolutionary processes rooted in fundamental biological requirements, including sustained proliferation and nutrient utilization. We hypothesize that the accumulation of mutations activates an evolutionary switch, enabling tumor cells to acquire an enhanced capacity for survival, adaptation, and growth at rates far exceeding typical evolutionary timescales. We propose the "evolutionary cascade hypothesis," a unifying framework that integrates these models into a coherent sequence. At its core lies the failure of DNA repair mechanisms, representing a critical transition in cancer progression. This shift marks the transition from an initial non-Darwinian, neutral phase to a Darwinian, more deterministic phase. CONCLUSIONS: As predictive models of tumor evolution advance through genomic big data and artificial intelligence-driven analysis, the future of cancer treatment may extend beyond targeting individual mutations to disrupting the underlying evolutionary mechanisms that sustain malignancy. This paradigm shift could redefine therapeutic strategies and ultimately improve patient outcomes.

Humans↗

An ozone budget for the UK: using measurements from the national ozone monitoring network; measured and modelled meteorological data, and a 'big-leaf' resistance analogy model of dry deposition.

Data from the UK national air-quality monitoring network are used to calculate an annual mass budget for ozone (O3) production and loss in the UK boundary layer during 1996. Monthly losses by dry deposition are quantified from 1 km x 1 km scale maps of O(3) concentration and O(3) deposition velocities based on a big-leaf resistance analogy. The quantity of O(3) deposition varies from approximately 50 Gg-O(3) month(-1) in the winter to over 200 Gg-O(3) month(-1) in the summer when vegetation is actively absorbing O(3). The net O(3) production or loss in the UK boundary layer is found by selecting days when the UK is receiving "clean" Atlantic air from the SW to NW. In these conditions, the difference in O(3) concentration observed at Mace Head and a rural site on the east coast of the UK indicates the net O(3) production or loss within the UK boundary layer. A simple box model is then used to convert the concentration difference into a mass. The final budget shows that for most of the year the UK is a net sink for O(3) (-25 to -800 Gg-O(3) month(-1)) with production only exceeding losses in the photochemically active summer months (+45 Gg-O(3) month(-1)).

Air Pollutants↗

Streamlining large-scale genomic data management: Insights from the UK Biobank whole-genome sequencing data.

Biobank-scale whole-genome sequencing (WGS) studies are increasingly pivotal in unraveling the genetic bases of diverse health outcomes. However, managing and analyzing these datasets' sheer volume and complexity presents significant challenges. We highlight the annotated genomic data structure (aGDS) format, substantially reducing the WGS data file size while enabling seamless integration of genomic and functional information for comprehensive WGS analyses. The aGDS format yielded 23 chromosome-specific files for the UK Biobank 500k WGS dataset, occupying only 1.10 tebibytes of storage. We develop the vcf2agds toolkit that streamlines the conversion of WGS data from VCF to aGDS format. Additionally, the STAARpipeline equipped with the aGDS files enabled scalable, comprehensive, and functionally informed WGS analysis, facilitating the detection of common and rare coding and noncoding phenotype-genotype associations. Overall, the vcf2agds toolkit and STAARpipeline provide a streamlined solution that facilitates efficient data management and analysis of biobank-scale WGS data across hundreds of thousands of samples.

Humans↗

Analyses of Digman's child-personality data: derivation of Big-Five factor scores from each of six samples.

One of the world's richest collections of teacher descriptions of elementary-school children was obtained by John M. Digman from 1959 to 1967 in schools on two Hawaiian islands. In six phases of data collection, 88 teachers described 2,572 of their students, using one of five different sets of personality variables. The present report provides findings from new analyses of these important data, which have never before been analyzed in a comprehensive manner. When factors developed from carefully selected markers of the Big-Five factor structure were compared to those based on the total set of variables in each sample, the congruence between both types of factors was quite high. Attempts to extend the structure to 6 and 7 factors revealed no other broad factors beyond the Big Five in any of the 6 samples. These robust findings provide significant new evidence for the structure of teacher-based assessments of child personality attributes.

Affect↗

Distribution and disappearance of exogenous [125I] big renin in the newborn puppy.

The distribution and metabolism of infused exogenous [125I] inactive plasma big renin, molecular weight 56,000 was studied in five newborn puppies. The animals were sacrificed and the organs removed and studied by chromatography, along with periodic blood samples taken during the 120-min study, for evidence of conversion of high-molecular-weight renin to low-molecular-weight renin. The decay curve suggested an initial rapid distribution (alpha) phase of 10 +/- 1.5 min followed by a slower elimination (beta) phase of 40 +/- 4.6 min [125I]-Big renin was taken up by the red blood cell and released slowly. The liver, kidneys, and lungs had the highest % of [125I]-big renin at the termination of the study. There was no chromatographic evidence of a change in molecular weight of the [125I]-big renin. These data show that big renin has a two-compartment disappearance curve and that there is no evidence of conversion of high-molecular-weight renin to low-molecular-weight renin systemically or in the tissues of the newborn canine puppy.

Animals↗

A comparison of cold and acid activation of big renin and of inactive renin in normal plasma.

Normal human plasma contains "inactive renin," whose ability to generate angiotensin I increases after exposure to pH 3.3. Big renin is a partially inactive enzyme of larger molecular weight, which is also activated at pH 3.3, and is found of pregnant women, and in amniotic fluid, but not in normal plasma. We have compared the effects of acid exposure and storage at 4 and -4 C on normal plasma and plasma containing big renin. The concentration of inactive renin in normal plasma was approximately equal to that of normal active renin, and its activity increased slowly on prolonged standing at -4 but not 4 C. In contrast, the activity of big renin increased by 50% as early as 1-3 days at 4 C and increased even more quickly at -4 C. Acid treatment of plasma containing big renin caused 4-10 times greater increase in active renin than similar treatment of normal plasma. During gel filtration, both cold-activated and previously acidified big renin coeluted with unactivated big renin. These data indicate that big renin is highly susceptible to cold or acid activation and that such activation of big renin does not result in a detectable decrease in its molecular weight of 60,000 daltons. Furthermore, acid and cold seem to activate the same pool of inactive renin in normal plasma. Although both normal and big renin are stable for long periods below -20 C, a serious overestimate of plasma renin activity can occur if plasma is stored just above its freezing point before assay.

Amniotic Fluid↗