Search PubMedSearch

SEARCH · Search PubMed

Results for “python”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

196 records · Page 11Linked to original sources

A comparative study on glyoxalase II from vertebrata.

S-2-hydroxyacylglutathione hydrolase (glyoxalase II) from the liver of animals belonging to the various vertebrate classes (Oryctolagus cuniculus, Gallus gallus, Python molurus, Rana esculenta, Esox lucius) have been purified from 100,000 g supernatants of liver homogenates, using acetone fractionation and affinity chromatography. Subsequent comparative studies were concerned with some molecular and kinetic properties. Isoelectric focusing gave evidence for a single form of liver glyoxalase II in O. cuniculus, P. molurus and E. lucius, while the enzyme from G. gallus and R. esculenta showed respectively two and three forms with different pI values. All studied enzymes are basic proteins. The relative molecular mass values range from 18,000 to 23,000. The various glyoxalases II do not display markedly different Kn or Ki values. Their stability behavior at different temperatures is also quite similar.

Animals

Localization of immunoreactive synthetic atrial natriuretic factor (ANF) in the heart of various animal species.

The localization of two synthetic fragments of the C-terminal portion of atrial natriuretic factor: Arg 101-Tyr 126 which displays full biological activity and Leu 94-Arg 109 which is completely devoid of biological activity, has been investigated by immunohisto- and immunocytochemical methods in the heart of mammals (rat, mouse, guinea pig, hamster, rabbit, cat, dog, man) and nonmammalian vertebrates toad (Bufo marinus), frog (Rana catesbeiana), fish (Cyprinus carpio, Puntius schwanenfeldi, Cichlosoma biocellatum, Carrasius auratus), snake (Python reticulatus) and hen. Antibodies against the synthetic fragments of ANF were raised in rabbits and used either for immunofluorescence (Coons' technique), immunohistochemistry (unlabeled antibody technique) or immunocytochemistry (protein A-gold technique). Results obtained by immunofluorescence and by the unlabeled antibody technique were similar: antibodies against Arg 101-Tyr 126 ANF allowed visualization of granulated cardiocytes in the atria of all mammals. While the reaction was very strong in rat and mouse, it was less so in the rabbit and very weak in all other species studied including man. Antibodies against Leu 94-Arg 109 ANF produced a reaction only in the rat and mouse. In nonmammalian vertebrates, the reaction was always much stronger in atria than ventricles of all species with both antibodies.

Animals

Comprehensive in silico genomics analysis of global trends and host-specific emergence of aminoglycoside resistance in Staphylococcus aureus: a One-Health perspective.

BACKGROUND: Aminoglycosides remain clinically valuable against Staphylococcus aureus. Aminoglycoside resistance in S. aureus represents a critical One Health concern and is primarily driven by aminoglycoside-modifying enzymes (AMEs), which are frequently plasmid-encoded. Although regional studies have provided valuable insights, the global epidemiology of aminoglycoside resistance determinants remains poorly characterized because comprehensive data integrating human, animal, and environmental reservoirs are still lacking. This study addresses this gap by analyzing over 110,000 S. aureus genomes (2000-2025) to map the global resistome, quantify temporal and host-specific trends, and assess the association between genetic determinants and phenotypic resistance. METHODS: We performed a retrospective One Health meta-analysis of 110,309 S. aureus genomes collected between 2000 and 2025 from 128 countries. Genomes were quality-filtered and aminoglycoside resistance determinants were identified using NCBI AMRFinderPlus (v4.0.23). Multilocus sequence typing and host-source harmonization (Human, Animal, Environment, Unknown) enabled clonal and reservoir stratification. Temporal trends in gene prevalence and resistance burden were modeled with robust regression. Geographic and host-associated structuring of key genes was assessed via &#x3c7;2 and enrichment tests. Machine-learning models (elastic-net, random forests, XGBoost) were benchmarked for minimum inhibitory concentration (MIC) prediction via nested cross-validation, with performance evaluated by mean absolute error, RMSE, and SHAP-based feature importance. All analyses were conducted in R and Python using publicly available, de-identified genomic data. RESULTS: Aminoglycoside resistance-associated genes were dominated by modifying enzyme determinants, with ant(6)-Ia, ant(9)-Ia, aph(3')-IIIa, sat4, aadD1, and aac(6')-Ie/aph(2'')-Ia occurring in 14-22% of isolates worldwide. Temporal analysis revealed significant declines in several major determinants, most notably ant(9)-Ia (-2.22 percentage points per year, p&#x2009;<&#x2009;0.001), whereas apmA exhibited a non-significant decreasing trend in animal isolates. Host structuring was marked: human clinical isolates concentrated common determinants, while animal and environmental isolates harbored rare alleles (apmA, spw, str, spd). Geographic mapping confirmed near-universal distribution of common genes but focal restriction of rare ones. Publicly available phenotypic data indicated strong activity of amikacin, whereas gentamicin showed a distinct resistant subpopulation that closely corresponded with AME gene carriage. Genotype-phenotype analyses demonstrated strong concordance, with gene-rich complements predicting resistant MIC strata and absence of determinants predicting susceptibility. Analysis across different gene classes revealed frequent co-occurrence of aminoglycoside resistance genes with determinants from other classes, such as mecA, blaZ, and MLS_B, embedding them within multidrug-resistant (MDR) genomic contexts. CONCLUSION: Over 25&#xa0;years, the prevalence of aminoglycoside resistance-associated genes in S. aureus has declined for several common determinants, while rare veterinary-linked alleles are emerging in animal isolates. Strong genotype-phenotype concordance supports genomic prediction for gentamicin and amikacin, where MIC data are available, although phenotypic confirmation remains essential. The frequent co-occurrence of aminoglycoside resistance genes with other antimicrobial resistance determinants indicates their integration within co-occurrence patterns of MDR genes, defined here as clusters of co-occurring resistance genes often carried on shared mobile genetic elements. These patterns highlight the need for integrated One Health surveillance combining clinical, veterinary, and environmental monitoring with plasmid-context resolution to anticipate emerging threats.

Aminoglycosides

Anticancer drug response prediction integrating multi-omics pathway-based difference features and multiple deep learning techniques.

Individualized prediction of cancer drug sensitivity is of vital importance in precision medicine. While numerous predictive methodologies for cancer drug response have been proposed, the precise prediction of an individual patient's response to drug and a thorough understanding of differences in drug responses among individuals continue to pose significant challenges. This study introduced a deep learning model PASO, which integrated transformer encoder, multi-scale convolutional networks and attention mechanisms to predict the sensitivity of cell lines to anticancer drugs, based on the omics data of cell lines and the SMILES representations of drug molecules. First, we use statistical methods to compute the differences in gene expression, gene mutation, and gene copy number variations between within and outside biological pathways, and utilized these pathway difference values as cell line features, combined with the drugs' SMILES chemical structure information as inputs to the model. Then the model integrates various deep learning technologies multi-scale convolutional networks and transformer encoder to extract the properties of drug molecules from different perspectives, while an attention network is devoted to learning complex interactions between the omics features of cell lines and the aforementioned properties of drug molecules. Finally, a multilayer perceptron (MLP) outputs the final predictions of drug response. Our model exhibits higher accuracy in predicting the sensitivity to anticancer drugs comparing with other methods proposed recently. It is found that PARP inhibitors, and Topoisomerase I inhibitors were particularly sensitive to SCLC when analyzing the drug response predictions for lung cancer cell lines. Additionally, the model is capable of highlighting biological pathways related to cancer and accurately capturing critical parts of the drug's chemical structure. We also validated the model's clinical utility using clinical data from The Cancer Genome Atlas. In summary, the PASO model suggests potential as a robust support in individualized cancer treatment. Our methods are implemented in Python and are freely available from GitHub (https://github.com/queryang/PASO).

Deep Learning

'PePApipe': A complete bioinformatics analysis pipeline for African Swine Fever Virus genome.

African Swine Fever Virus (ASFV) is of high concern in porcine livestock across the world due to both the high mortality rates and the trade restrictions imposed on affected regions. The viral genome is large and complex, and genomic analysis is essential for tracing its origin and evolution. Although several bioinformatics tools exist for genome assembly and analysis, no single platform integrates all necessary steps in an accessible and systematic way. In this study the authors developed 'PePApipe', a custom-built, user-friendly pipeline that enables rapid, complete, and efficient ASFV genome analysis. It is specifically designed for laboratory professionals with limited bioinformatics experience, requiring only basic command-line knowledge. Starting from raw sequencing data, PePApipe integrates thirteen software tools into one automated workflow, covering quality control and pre-processing of raw reads, de novo genome assembly and variant calling. Programmed in Python, it can be executed locally through bash scripts, or using a Slurm protocol for batch processing of multiple samples. The main outputs are the ASFV consensus genome sequence and a file listing its putative variants compared to the selected reference genome. PePApipe classifies generated files into structured folders and produces intermediate files that can be used as inputs for further or parallel analyses; users can also enable or disable specific steps in each particular case. This pipeline is adaptable and complementary to downstream steps such as viral genome annotation or genome visualization. By consolidating all stages of viral genome analysis into a single automated workflow, PePApipe reduces the likelihood of user error, and enhances reproducibility and efficiency. This user-friendly pipeline facilitates the transition from sequencing to assembly and downstream analysis of viral genomes, ensuring a fast and reliable response to molecular analysis demands. Finally, the pipeline can be easily adapted to the study of other viral species, expanding its application in infectious diseases surveillance.

African Swine Fever Virus

REAPER: a project-centric workflow layer for comparative repeatome analysis.

INTRODUCTION: Repeatome characterization from short-read sequencing data is widely performed using RepeatExplorer2/TAREAN. However, long-lived multisample projects and explicit comparative designs are often executed as ad hoc command sequences that are hard to version, rerun, and monitor on shared compute environments - a gap that motivates a project-centric workflow layer for repeatome analysis. METHODS: We present REAPER (Repeatome Extended Analysis Pipeline-Execution and Reporting), a project-centric workflow layer that couples a modular Snakemake pipeline with a Python project manager to enforce a stable on-disk layout and configuration-driven execution for single-sample and comparative repeatome analyses. REAPER does not implement a new repeat-discovery algorithm; it is an orchestration layer, and biological accuracy for clustering and satellite calling depends on the underlying RepeatExplorer2/TAREAN and satMiner methods it coordinates. REAPER standardizes: Read QC Deterministic subsampling and preparation RepeatExplorer2/TAREAN execution via seqclust, with satMiner-inspired iterative assembly Post-TAREAN BLAST-based annotation against curated repeat collections (optionally including taxon-scoped NCBI-derived resources with freshness checks) Optional graph-based comparative reports The pipeline makes comparative read allocation, prefix policy, and analysis-ready tables explicit; caching supports incremental reruns and structured logs support monitoring. Performance was assessed using a Triticeae short-read dataset (five samples), with rule-level logging of runtime and memory across pipeline stages. RESULTS: Rule-level performance logs show that graph-based clustering dominates runtime and memory, while QC and preparation steps are lightweight by comparison. Graph-report annotations for the Triticeae project additionally link high-ranking clusters to established repeat markers - including pTa794- and pSc119-class entries in curated databases. DISCUSSION: These findings illustrate biologically interpretable outputs (recovery of known Triticeae repeat markers) alongside quantitative performance metrics (identification of graph-based clustering as the dominant computational cost). By making comparative read allocation, prefix policy, and analysis-ready tables explicit - and by supporting caching and structured logging - REAPER supports reproducible comparative repeatome analysis in evolving multisample projects. As an orchestration layer rather than a discovery algorithm, REAPER's contribution lies in reproducibility, monitorability, and comparative-analysis infrastructure, with biological accuracy remaining contingent on the underlying RepeatExplorer2/TAREAN and satMiner methods.

TAREAN

BioEMMA: Automated Generation of Model-Specific Escher-Compatible Maps from KEGG Pathways.

Genome-scale metabolic models are widely used to investigate cellular metabolism, but their interpretation and comparison are limited by the lack of reproducible pathway-level visualizations with a common spatial organization. This study presents BioEMMA, a Python-based tool for the automated generation of model-specific metabolic pathway maps in the Escher JSON format using coordinate information from curated KEGG pathway maps. BioEMMA parses KGML files, map reaction and metabolite identifiers to model database namespaces, filters pathway elements according to an input SBML model, adds non-primary metabolites, reconstructs Escher-compatible layouts, and supports flux visualization. The tool was integrated into a reproducible BioUML workflow for metabolic model reconstruction. BioEMMA was evaluated using the e_coli_core model and the KEGG glycolysis/gluconeogenesis pathway while generating a model-specific map with overlaid FBA fluxes. It was then applied to compare E. coli reconstructions generated by gapseq, ModelSEEDpy, and Reconstructor across three central carbon metabolism pathways. To broaden the evaluation, BioEMMA was applied using 87 prokaryotic BiGG models and three eukaryotic models. The analysis revealed pathway-specific differences in reaction coverage, shared and model-specific reactions, and predicted flux activity. BioEMMA therefore provides a reproducible framework for pathway-level visualization and comparison of genome-scale metabolic reconstructions within a common spatial coordinate system.

Escher maps

A Practical Workflow for Spatial Transcriptomics Data Analysis: From Data Acquisition to Advanced Analyses.

Spatial transcriptomics (ST) profiles genome-wide gene expression while preserving the two-dimensional spatial context of mRNA molecules within tissue sections, enabling studies of tissue architecture and microenvironment-associated biology. However, ST analysis remains challenging because data import, quality control, integration, deconvolution, spatial statistics, and visualization often require multiple software environments and reproducible parameter choices. This protocol presents a practical computational workflow for public ST datasets in R, beginning with data acquisition and software setup and proceeding through Seurat-based data loading, quality control, normalization, multi-sample integration, clustering, and spatially variable gene analysis. The workflow then applies complementary deconvolution strategies, including reference-guided SPOTlight analysis and unsupervised STdeconvolve topic modeling, followed by Giotto-based spatial cell-cell communication analysis and interactive region-of-interest (ROI) selection using a custom Python Dash application. By emphasizing script-based execution, explicit parameter rationales, expected outputs, and troubleshooting checkpoints, the protocol provides an adaptable framework for standard array-based ST datasets and related platforms after dataset- and platform-specific parameter evaluation.

Spatial Transcriptomics

POISE: Spectral Inference of Parent-of-Origin Effects in Unlabeled Genomic Data.

MOTIVATION: Parent of Origin Effects (POEs), where the effect of an an allele on a phenotype differs based on maternal or paternal inheritance implicated in growth, metabolism, and neurodevelopment. Traditional tests for POEs require family data to determine parental origins of transmitted alleles. Given that such studies are expensive and time consuming compared to genome-wide association studies (GWAS), tests that function absent inheritance information are highly desirable. We develop a method, based on community detection from machine learning, that infers POEs via a spectral decomposition, obtains confidence intervals via a non-parametric bootstrap, and safeguards against confounding by non POE sources of variation. We refer to our method as Parent of Origin Inference via Spectral Estimation (POISE). RESULTS: We demonstrate that POISE is well-calibrated under both Gaussian and heavy-tailed noise in simulation studies, with improved robustness to true POEs compared to existing covariance-based tests. POISE provides per-trait effect estimates with bias-corrected bootstrap confidence intervals and incorporates an information-theoretic minimum detectable effect size that filters unreliable estimates, conferring robustness to covariance-deflating variance QTL. We then apply POISE to GWAS data from the UK Biobank using BMI, LDL cholesterol, and HDL cholesterol. POISE recovers established POE loci and identifies 134 additional variants at genes implicated in lipid metabolism, immune regulation, and growth. AVAILABILITY AND IMPLEMENTATION: The code for this method in Python is available at https://github.com/bystrogenomics/POISE.

Community Detection

Phasis: a software tool for register-resolved discovery of plant phased small RNA loci.

Plant PHAS locus discovery remains challenging because phasiRNA-producing loci must be distinguished from other sRNA-producing regions with high abundance or apparent periodicity. This problem is especially acute for reproductive 24-PHAS loci, which occur within genomes that also produce abundant 24-nt siRNAs from nonPHAS regions. We present Phasis, an open-source Python software tool for plant PHAS-locus discovery from small RNA sequencing data. Phasis combines statistical evidence for phased accumulation with locus-level features and a Register-Resolved Locus Interpretation Layer that evaluates whether candidate loci show coherent phased architecture. Across diverse plant datasets, Phasis recovered validated or annotated 21- and 24-PHAS loci with a strong balance between call-level precision and reference-locus recall, and generally outperformed PhaseTank and ShortStack in matched benchmark analyses. The register-resolved interpretation layer reduced unsupported calls by separating coherent phased loci from ambiguous sRNA-producing regions. In maize dcl5 mutant libraries, Phasis showed strong depletion of 24-PHAS recovery, supporting DCL5-dependent recovery of reproductive 24-PHAS signal. Together, these results support Phasis as a biologically interpretable tool for large-scale discovery of plant DCL-dependent phasiRNA loci.

bioinformatics

Review of sarcocystosis in Malaysia.

Sarcocystis is a tissue coccidian with an obligatory two-host life cycle. The sexual generations of gametogony and sporogony occur in the lamina propria of the small intestine of definitive hosts which shed infective sporocysts in their stools and present with intestinal sarcocystosis. Asexual multiplication occurs in the skeletal and cardiac muscles of intermediate hosts which harbor Sarcocystis cysts in their muscles and present with muscular sarcocystosis. In Malaysia, Sarcocystis cysts have been reported from many domestic and wild animals, including domestic and field rats, moonrats, bandicoots, slow loris, buffalo, and monkey, and man. The known definitive hosts for some species of Sarcocystis are the domestic cat, dog and the reticulated python. Human muscular sarcocystosis in Malaysia is a zoonotic infection acquired by contamination of food or drink with sporocysts shed by definitive hosts. The cysts reported in human muscle resembled those seen in the moonrat, Echinosorex gymnurus, and the long-tailed monkey, Macaca fascicularis. While human intestinal sarcocystosis has not been reported in Malaysia so far, it can be assumed that such cases may not be infrequent in view of the occurrence of Sarcocystis cysts in meat animals, such as buffalo. The overall seroprevalence of 19.8% reported among the main racial groups in Malaysia indicates that sarcocystosis (both the intestinal and muscular forms) may be emerging as a significant food-borne zoonotic infection in the country.

Animals

Sero-epidemiologic investigations on brucellosis in the states of Uttar Pradesh (U.P.) and Delhi (India).

Sero-prevalence of brucellosis in man and animals was studied during the years 1976 and 1977. Samples were collected from Hospitals/slaughter houses/livestock farms located in Delhi and different districts of Uttar Pradesh (U.P.). The sera samples tested were from 1685 men, 1607 goats, 438 sheep, 244 pigs, 361 cattle, 551 buffalos, 50 dogs, 318 equine and 43 free living animals. The percentage of seropositivity, excluding doubtful ones, was recorded as: man 0.89, goat 5.53, sheep 3.42, pigs 15.98, cattle 6.37 buffalo 4.9 and equine 12.89. Additionally an evidence of agglutinins was also detected in a python serum sample. It was observed that occupation, age, sex and season had a bearing on the prevalence of the disease.

Animals

Host specificity and host range of the genus Sarcocystis in three snake-rodent life cycles.

Three Sarcocystis species with snake-rodent life cycles were studied for their host range and host specificity in systematically related intermediate and definitive hosts. While S. singaporensis and S. villivillosi developed only in murids closely related to the genus Rattus, a third Sarcocystis species (natural definitive host: Bitis nasicornis - syn. Isospora dirumpens ?) showed a broad intermediate host range. This species was found to use several Bitis species as definitive hosts. S. singaporensis and S. villivillosi in contrary produced sporocysts only in three Python species and in Aspidites melanocephalus .

Animals

Practicing Data Science in Interactive Notebooks.

The Jupyter Notebook is a platform for interactive computing that displays code and results in the same browser, making it valuable for teaching, prototyping, data analysis, and collaboration. Its explicit&#xa0;and transparent structure greatly reproducibility&#xa0;while&#xa0;its backend server supports flexible deployment. In the past few years, Jupyter notebooks and similar tools have become increasingly popular. In this chapter, we will review key aspects of data analysis in a cloud environment and demonstrate common tasks for analyzing metabolomics data using template notebooks. This is an accompaniment to the basic bioinformatics tools and essential data science toolkit introduced in the first edition.

Software

CNV-Finder: Streamlining Copy Number Variation Discovery.

Copy Number Variations (CNVs) play pivotal roles in the etiology of complex diseases and are variable across diverse populations. Understanding the association between CNVs and disease susceptibility is significant in disease genetics research and often requires analysis of large sample sizes. One of the most cost-effective and scalable methods for detecting CNVs is based on normalized signal intensity values, such as Log R Ratio (LRR) and B Allele Frequency (BAF), from Illumina genotyping arrays. In this study, we present CNV-Finder, a novel pipeline integrating deep learning techniques on array data, specifically a Long Short-Term Memory (LSTM) network, to expedite the large-scale identification of CNVs within predefined genomic regions. This facilitates efficient prioritization of samples for time-consuming or costly subsequent analyses such as Multiplex Ligation-dependent Probe Amplification (MLPA), short-read, and long-read whole genome sequencing. We incorporate four genes to establish our methods-Parkin (PRKN), Leucine Rich Repeat And Ig Domain Containing 2 (LINGO2), Microtubule Associated Protein Tau (MAPT), and alpha-Synuclein (SNCA)-which may be relevant to neurological diseases such as Alzheimer's disease (AD), Parkinson's disease (PD), Progressive Supranuclear Palsy (PSP), or related disorders such as essential tremor (ET). By training our models on expert-annotated samples and validating them across diverse cohorts, including those from the Global Parkinson's Genetics Program (GP2) and additional dementia-specific databases, we demonstrate the efficacy of CNV-Finder in accurately detecting deletions and duplications. Our pipeline outputs app-compatible files for visualization within CNV-Finder's interactive web application. This interface enables researchers to review predictions and filter displayed samples by model prediction values, LRR range, and variant count in order to explore or confirm results. Our pipeline integrates this human feedback to enhance model performance and reduce false positive rates. Through a series of comprehensive analyses and validations using visual inspection, MLPA, short-read, and long-read sequencing data, we demonstrate the robustness and adaptability of CNV-Finder in identifying CNVs with regions of varied size, probe density, and noise. Our findings highlight the significance of contextual understanding and human expertise in enhancing the precision of CNV identification, particularly in complex genomic regions like 17q21.31. The CNV-Finder pipeline is a scalable, publicly available resource for the scientific community, available on GitHub (https://github.com/GP2code/CNV-Finder; DOI 10.5281/zenodo.14182563). CNV-Finder not only expedites accurate candidate identification but also significantly reduces the manual workload for researchers, enabling future targeted validation and downstream analyses in regions or phenotypes of interest.

Copy Number Variation (CNV)