Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “network analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 739 records · Page 41Linked to original sources

Exploring similarity between peer educators and their contacts and AIDS-protective behaviours in reproductive health programmes for adolescents and young adults in Ghana.

This analysis explores the similarity between peer educators and their contacts. To examine interpersonal communication in the context of peer education, this study tested a new approach using multiple semi-structured interviews and network analysis to collect data from 106 peer educators and 526 of their contacts. These evaluation activities were conducted at three sites in Ghana during April 1998, in peri-urban and rural locations, and in in-school and out-of-school targeted settings. It was found that in their peer counselling and peer promotion activities peer educators tend to reach people who are like themselves (53% within 2 years of age, 59% same sex, 70% same ethnicity, and 65% same school status) however, this trend is not uniform among all youth and varies by demographic characteristics and their cultural environment. By examining the social networks of peer educators, it is possible to gain a better understanding of the process of peer education counselling in the context in which it occurs. The study also shows that controlling for other factors, contacts of peer educators who are highly similar regarding age, sex, ethnicity, and school status, are 1.74 times more likely (95% CI: 1.18, 2.56) to have done something to protect themselves from AIDS in the past three months. The results have relevance for programme managers and planners, researchers, and international agencies serving youth.

Acquired Immunodeficiency Syndrome↗

Systematic analysis of yeast strains with possible defects in lipid metabolism.

Lipids are essential components of all living cells because they are obligate components of biological membranes, and serve as energy reserves and second messengers. Many but not all genes encoding enzymes involved in fatty acid, phospholipid, sterol or sphingolipid biosynthesis of the yeast Saccharomyces cerevisiae have been cloned and gene products have been functionally characterized. Less information is available about genes and gene products governing the transport of lipids between organelles and within membranes or the turnover and degradation of complex lipids. To obtain more insight into lipid metabolism, regulation of lipid biosynthesis and the role of lipids in organellar membranes, a group of five European laboratories established methods suitable to screen for novel genes of the yeast Saccharomyces cerevisiae involved in these processes. These investigations were performed within EUROFAN (European Function Analysis Network), a European initiative to identify the functions of unassigned open reading frames that had been detected during the Yeast Genome Sequencing Project. First, the methods required for the complete lipid analysis of yeast cells based on chromatographic techniques were established and standardized. The reliability of these methods was demonstrated using tester strains with established defects in lipid metabolism. During these investigations it was demonstrated that different wild-type strains, among them FY1679, CEN.PK2-1C and W303, exhibit marked differences in lipid content and lipid composition. Second, several candidate genes which were assumed to encode proteins involved in lipid metabolism were selected, based on their homology to genes of known function. Finally, lipid composition of mutant strains deleted of the respective open reading frames was determined. For some genes we found evidence suggesting a possible role in lipid metabolism.

Antifungal Agents↗

Biological selection criteria for radical prostatectomy.

Tumors clinically confined to the prostate gland (T1-2) are heterogeneous with respect to pathological staging and outcome after definitive radical surgery (radical prostatectomy). The preoperative prognostic factors that could predict pathological stage and outcome of individual patients with clinically localized prostate cancer are reviewed. New preoperative factors have been identified by histological analysis of needle biopsy prostate specimens in addition to Gleason grading score, serum markers (PSA), and clinical staging. These factors are related to tumor volume, zonal origin of the tumor, and spread into the gland and surrounding tissues. Other biological factors are identified by molecular and immunohistochemical analysis (neuroendocrine differentiation, DNA content, microvessel density, and perineural invasion). Biomolecular factors can also be assessed preoperatively on serum samples (free/total PSA ratio, PSA RT-PCR). Although only a few of these factors have a role in predicting treatment failure and/or disease recurrence, the neural network analysis seems to be the most important tool for identifying patients with more aggressive disease. A combination of these new factors, also using neural networks, could be relevant in the preoperative management of patients with prostate cancer to identify those with confined disease and to select those suitable for a "nerve sparing radical prostatectomy" to preserve sexual function and to achieve definitive cancer control.

Biopsy↗

Health outcomes associated with various antihypertensive therapies used as first-line agents: a network meta-analysis.

CONTEXT: Establishing relative benefit or harm from specific antihypertensive agents is limited by the complex array of studies that compare treatments. Network meta-analysis combines direct and indirect evidence to better define risk or benefit. OBJECTIVE: To summarize the available clinical trial evidence concerning the safety and efficacy of various antihypertensive therapies used as first-line agents and evaluated in terms of major cardiovascular disease end points and all-cause mortality. DATA SOURCES AND STUDY SELECTION: We used previous meta-analyses, MEDLINE searches, and journal reviews from January 1995 through December 2002. We identified long-term randomized controlled trials that assessed major cardiovascular disease end points as an outcome. Eligible studies included both those with placebo-treated or untreated controls and those with actively treated controls. DATA EXTRACTION: Network meta-analysis was used to combine direct within-trial between-drug comparisons with indirect evidence from the other trials. The indirect comparisons, which preserve the within-trial randomized findings, were constructed from trials that had one treatment in common. DATA SYNTHESIS: Data were combined from 42 clinical trials that included 192 478 patients randomized to 7 major treatment strategies, including placebo. For all outcomes, low-dose diuretics were superior to placebo: coronary heart disease (CHD; RR, 0.79; 95% confidence interval [CI], 0.69-0.92); congestive heart failure (CHF; RR, 0.51; 95% CI, 0.42-0.62); stroke (RR, 0.71; 0.63-0.81); cardiovascular disease events (RR, 0.76; 95% CI, 0.69-0.83); cardiovascular disease mortality (RR, 0.81; 95% CI, 0.73-0.92); and total mortality (RR, 0.90; 95% CI, 0.84-0.96). None of the first-line treatment strategies-beta-blockers, angiotensin-converting enzyme (ACE) inhibitors, calcium channel blockers (CCBs), alpha-blockers, and angiotensin receptor blockers-was significantly better than low-dose diuretics for any outcome. Compared with CCBs, low-dose diuretics were associated with reduced risks of cardiovascular disease events (RR, 0.94; 95% CI, 0.89-1.00) and CHF (RR, 0.74; 95% CI, 0.67-0.81). Compared with ACE inhibitors, low-dose diuretics were associated with reduced risks of CHF (RR, 0.88; 95% CI, 0.80-0.96), cardiovascular disease events (RR, 0.94; 95% CI, 0.89-1.00), and stroke (RR, 0.86; 0.77-0.97). Compared with beta-blockers, low-dose diuretics were associated with a reduced risk of cardiovascular disease events (RR, 0.89; 95% CI, 0.80-0.98). Compared with alpha-blockers, low-dose diuretics were associated with reduced risks of CHF (RR, 0.51; 95% CI, 0.43-0.60) and cardiovascular disease events (RR, 0.84; 95% CI, 0.75-0.93). Blood pressure changes were similar between comparison treatments. CONCLUSIONS: Low-dose diuretics are the most effective first-line treatment for preventing the occurrence of cardiovascular disease morbidity and mortality. Clinical practice and treatment guidelines should reflect this evidence, and future trials should use low-dose diuretics as the standard for clinically useful comparisons.

Humans↗

Multisolutional clustering and quantization algorithm (MCQ).

We have developed a novel clustering and quantization algorithm that allows the user to create multiple one-to-one correspondences between the actual data and its transformed (clustered and quantized) values, based on the user's hypothesis regarding the nature of the classification task. The types of problems for which the algorithm can be beneficial are discussed. We report experiments employing simulated and real data that suggest the proposed algorithm may be useful in neural network analysis of various phenomena in medicine and biology.

Algorithms↗

Carboxylic acids: prediction of retention data from chromatographic and electrophoretic behaviours.

A review of the main results reached in the prediction of retention data of carboxylic acids, inferred by their chromatographic and electrophoretic behaviour, is presented. Attention has been focused on the main separation methods used in carboxylic acids analysis, that is ion-exclusion, anion-exchange, reversed-phase (RP) liquid chromatography and capillary electrophoresis. Papers proposing mechanistic models as well as chemometric and multilayer feed-forward neural network analysis of ion chromatography (IC) and RP chromatographic retention data were reviewed. Principal component analysis, PCA, sequential simplex method and simultaneous modelling of response surfaces through simple nonlinear models (not related to equilibria involved in retention) have been considered. Computer simulations for the prediction of retention data have also been discussed. A quick overlook on the prediction of capacity factors of analytes by less common determination methods such as thin-layer, gas chromatography and supercritical fluid chromatography has also been done.

Carboxylic Acids↗

The church family and kin: an older rural black woman's support network and preferences for care providers.

Although kin and church are considered premier support sources for rural elders, few scholars have undertaken descriptive studies to explore the nature of rural Black elders' support networks and their preferences for in-home service providers. In the case study described in this article, methods of support network analysis and descriptive phenomenology were used to analyze data from five lengthy, open-ended interviews with a 94-year-old rural Black woman. The various groups and individuals of her network are labeled in her words, the network's supportive functions are described, and preferences for providers are noted. In addition, the varying structures of her home care experience with the support network members are described. Her attempts to voice and exercise her preferences for in-home service providers are explained in terms of two contrasting processes: preference uptake and preference suppression. Based on these findings, implications for appraising the appropriateness of rural elders' in-home services are discussed.

Black or African American↗

Distinct spatial transcriptomic patterns of substantia Nigra in Parkinson disease and Parkinsonian subtype of multiple system atrophy.

To investigate transcriptomic signatures of Parkinson's disease (PD) and the Parkinsonian subtype of Multiple System Atrophy (MSA-P) in substantia nigra pars compacta (SNpc), we conducted transcriptome analysis using in-situ hybridization on paraffin-embedded SNpc tissues from post-mortem brains. The study included 2 MSA-P patients, 2 PD patients, and 2 healthy controls (HC), with 12 regions of interest (ROIs) selected from the dorsal to ventral and medial to lateral aspects of the SNpc. A total of 72 ROIs from 6 participants were analyzed, and differentially expressed genes (DEGs) were identified by comparing MSA-P, PD and HC groups. The MSA-P group showed 88 upregulated DEGs and 326 downregulated DEGs (adjusted &#x1d45d;<0.05) compared to HC. The downregulated DEGs were significantly enriched in pathways related to ribosomal translation, immune processes, mitochondrial function, and autophagy. Notably, the dorsomedial quadrant was uniquely linked to antigen presentation, while other quadrants showed downregulation of protein synthesis. The PD group exhibited 165 upregulated DEGs and 350 downregulated DEGs (adjusted &#x1d45d;<0.05) compared to HC, with downregulated DEGs associated with ribosomal translation, mitochondrial function, and the ubiquitin-proteasome system. In both MSA-P and PD, the upregulated DEGs were not associated with any pathways or biological process in gene enrichment analysis. In network propagation analysis, amyloid precursor protein was the most significant network hub among DEGs in both MSA-P and PD. Comparing the transcriptomic signatures of SNpc between MSA-P and PD, we found immune/inflammation, mitochondrial function and neural signaling related genes were significantly downregulated in MSA-P compared to PD. Overall, the transcriptomic signature of the SNpc in MSA-P and PD revealed overlapping but distinct features, including alterations in protein synthesis, immune processes, mitochondrial function, and protein degradation systems. Future studies with larger cohorts and functional validation are needed to further elucidate these findings.

Humans↗

Protein crystallization: virtual screening and optimization.

Advances in genomics have yielded entire genetic sequences for a variety of prokaryotic and eukaryotic organisms. This accumulating information has escalated the demands for three-dimensional protein structure determinations. As a result, high-throughput structural genomics has become a major international research focus. This effort has already led to several significant improvements in X-ray crystallographic and nuclear magnetic resonance methodologies. Crystallography is currently the major contributor to three-dimensional protein structure information. However, the production of soluble, purified protein and diffraction-quality crystals are clearly the major roadblocks preventing the realization of high-throughput structure determination. This paper discusses a novel approach that may improve the efficiency and success rate for protein crystallization. An automated nanodispensing system is used to rapidly prepare crystallization conditions using minimal sample. Proteins are subjected to an incomplete factorial screen (balanced parameter screen), thereby efficiently searching the entire "crystallization space" for suitable conditions. The screen conditions and scored experimental results are subsequently analyzed using a neural network algorithm to predict new conditions likely to yield improved crystals. Results based on a small number of proteins suggest that the combination of a balanced incomplete factorial screen and neural network analysis may provide an efficient method for producing diffraction-quality protein crystals.

Combinatorial Chemistry Techniques↗

Networks of persons with syphilis and at risk for syphilis in Louisiana: evidence of core transmitters.

BACKGROUND AND OBJECTIVES: Differences in sociodemographic attributes and healthcare access may explain differences in regional sexually transmitted disease rates but don't fully explain why syphilis persists disproportionately in certain populations. GOAL OF THIS STUDY: To understand the behavioral epidemiology of syphilis, we conducted a social network analysis of persons with syphilis and their contacts and developed and applied a definition of core transmitters. STUDY DESIGN: We interviewed 10 index persons with primary or secondary untreated syphilis and 80 of their named sexual and social contacts. RESULTS: Fourteen (16%) of 90 interviewed persons met the definition of core transmitters, 9 of whom had past or current syphilis. The other interviewed persons had only moderately risky behaviors. Seventy-eight (42%) of the network sexual contacts were connected directly or indirectly to a core transmitter. CONCLUSION: This analysis suggests that syphilis transmission is maintained by a community with a small percentage of high-risk persons centrally placed amidst a larger group with moderately risky behavior.

Adult↗

Fixed-mass multifractal analysis of river networks and braided channels.

A fixed-mass multifractal (FMA) analysis was used to investigate natural river networks and braided channels. In particular, while the study of natural river networks was performed with fixed-size algorithms (FSAs) in the past, the analysis of natural braided channels was not pursued before to our knowledge. Results showed the multifractal and non-plane-filling nature of all the digitalized data sets. Analysis of the digitalization step (constant or not) was performed and showed that it does not exert a strong influence on the assessed values of the Lipschitz-Hölder exponents and the support dimensions, even if a constant step permits better reconstruction of the right sides of the spectra, for negative moment orders of probabilities. The FMA approach presented two improvements with respect to the FSA one, in terms of oscillations of the scaling curves for negative moment orders of probabilities and of error bars. A more precise assessment of the multifractal spectra is of great importance in the development of multifractal models for the simulation of flood hydrographs.

Journal Article↗

Measurement of the adhesive force of fine particles on tablet surfaces and method of their removal.

The adhesion force of fine particles on the surface of tablets was measured by a centrifugal force and impact separation method. A Finededuster (FDD) was employed to remove fine particles from the tablet surface. The centrifugal force and impact separation method was suggested to be effective for measuring the adhesive forces between particles and the tablet surface, and effective disjoining force in the FDD could be estimated by comparison of the results obtained using these two methods. The FDD showed high removal efficiency regardless of how many tablets were processed at the same time. In either of these methods, critical particle size was about 10-20 microns, and larger particles were removed more efficiently. This critical particle size was similar to that observed for other mechanical properties of powders, such as angle of repose and flowability. We simulated particle residual percentage under various operation conditions by ANN (artificial neural network) analysis and multiple regression analysis. This simulation enabled us to predict how the efficiency of particle removal is affected by the interaction of the experimental and material factors.

Administration, Oral↗

A Risk Score for Polycystic Ovary Syndrome Based on Meta-Analysis and Machine Learning of Gut Microbiota Signatures.

Polycystic Ovary Syndrome (PCOS) is a prevalent endocrine and metabolic disorder among reproductive-age women, in which emerging evidence suggests a substantial role played by the gut microbiota. To comprehensively evaluate gut microbiota alterations in PCOS and identify microbial biomarkers through integrated analysis, a systematic search of PubMed, Web of Science, and Embase was conducted for studies employing 16S rRNA gene sequencing of fecal samples from PCOS cohorts. Ten eligible PCOS cohorts, comprising 858 individuals, were included in the study, from which a risk score was derived using a 20-gene gut microbial signature associated with PCOS. Meta-analysis at the genus level identified that Subdoligranulum, NK4A214_group, and Collinsella significantly decreased, and Bacteroides increased in PCOS across multiple cohorts. Machine learning analysis identified a 20-genus microbial signature using the least absolute shrinkage and selection operator (LASSO) method, which was used to construct a risk score with an AUC of 0.835 in diagnosis prediction. Network analysis further identified Negativibacillus and Lachnospiraceae_UCG_010 as potential driver microbes in PCOS. The analysis in this study highlights key alterations in the gut microbiota across PCOS cohorts. The identified gut microbial signature and derived LASSO-based risk model offer novel insights and a potential tool for PCOS diagnosis.

Polycystic Ovary Syndrome↗

CCT2 defines a highly cisplatin-resistant and poor-prognosis subtype of lung adenocarcinoma.

Cisplatin-based chemotherapy is a standard treatment for lung adenocarcinoma (LUAD), yet acquired cisplatin resistance remains a marked cause of treatment failure. The molecular mechanisms driving cisplatin resistance in LUAD have not been fully elucidated. The present study integrated bulk transcriptomic data, genomic mutation profiles and single-cell RNA sequencing data to systematically investigate cisplatin resistance in LUAD. Resistance-associated genes were identified through differential expression, survival analysis and database integration. Unsupervised clustering was used to define cisplatin resistance-associated subtypes. Functional characteristics were explored using pathway enrichment, immune infiltration, tumor mutation burden and weighted gene co-expression network analysis. A machine learning framework incorporating 101 algorithms was applied to identify key genes and construct a prognostic model. Single-cell analyses and in vitro experiments were performed to validate the biological role of the core gene. Molecular docking and molecular dynamics simulations were conducted to identify potential therapeutic compounds. A total of two molecular subtypes with distinct cisplatin resistance levels and prognostic outcomes were identified. The high-resistance subtype exhibited enhanced cell cycle activity, DNA repair signaling and immune heterogeneity. Machine learning analysis revealed a five-gene signature, with chaperonin-containing TCP1 subunit 2 (CCT2) emerging as a key regulator of cisplatin resistance. Single-cell analyses showed that CCT2 was predominantly enriched in resistant epithelial cell subpopulations. Functional experiments demonstrated that CCT2 knockdown significantly inhibited cell proliferation and enhanced cisplatin sensitivity in LUAD cell lines. A number of candidate compounds targeting CCT2 exhibited stable binding in silico. The present findings identified CCT2 as a key mediator of cisplatin resistance in LUAD and provided potential therapeutic strategies to overcome chemotherapy resistance.

chaperonin-containing TCP-1 subunit 2↗

Identification of large-scale networks in the brain using fMRI.

Cognition is thought to result from interactions within large-scale networks of brain regions. Here, we propose a method to identify these large-scale networks using functional magnetic resonance imaging (fMRI). Regions belonging to such networks are defined as sets of strongly interacting regions, each of which showing a homogeneous temporal activity. Our method of large-scale network identification (LSNI) proceeds by first detecting functionally homogeneous regions. The networks of functional interconnections are then found by comparing the correlations among these regions against a model of the correlations in the noise. To test the LSNI method, we first evaluated its specificity and sensitivity on synthetic data sets. Then, the method was applied to four real data sets with a block-designed motor task. The LSNI method correctly recovered the regions whose temporal activity was locked to the stimulus. In addition, it detected two other main networks highly reproducible across subjects, whose activity was dominated by slow fluctuations (0-0.1 Hz). One was located in medial and dorsal regions, and mostly overlapped the "default" network of the brain at rest [Greicius, M.D., Krasnow, B., Reiss, A.L., Menon, V., 2003. Functional connectivity in the resting brain: a network analysis of the default mode hypothesis. Proceedings of the National Academy of Sciences of the U.S.A. 100, 253-258]; the other was composed of lateral frontal and posterior parietal regions. The LSNI method we propose allows to detect in an exploratory and systematic way all the regions and large-scale networks activated in the working brain.

Adult↗

Genomic mapping of diabetic kidney disease biomarkers and identification of potential inhibitors through virtual screening.

BACKGROUND: Diabetic kidney disease (DKD) is a common and serious complication of diabetes mellitus, marked by a multifactorial pathogenesis and the absence of sensitive diagnostic biomarkers. Identifying novel molecular targets and therapeutic options is essential to improve early diagnosis and treatment outcomes. METHODS: To uncover potential biomarkers and therapeutic candidates, we performed an integrated genomic analysis using microarray and RNA-seq datasets from the Gene Expression Omnibus (GEO) and Sequence Read Archive (SRA) databases. Differentially expressed genes (DEGs) were identified and subjected to protein-protein interaction (PPI) network analysis. Key genes were further explored through virtual screening of an FDA-approved compound library using molecular docking techniques. Drug-likeness was assessed via Lipinski's rule of five. RESULTS: A total of 40 DEGs were identified, among which ISCU (downregulated; involved in iron-sulfur cluster biogenesis) and AP1S2 (upregulated; associated with vesicular trafficking) emerged as potential biomarkers. PPI analysis revealed their involvement in critical DKD-related pathways, such as extracellular matrix remodeling and oxidative stress. Virtual screening identified six FDA-approved compounds with high binding affinity (&#x2264;-7.96 kcal/mol) to ISCU, notably ZINC000001576020, all of which complied with Lipinski's rule. CONCLUSIONS: This in-silico study nominates ISCU and AP1S2 as candidate diagnostic biomarkers for DKD and identifies computationally prioritized inhibitors targeting ISCU. These findings require experimental validation but provide a molecular framework for precision diagnosis and therapeutic development. These findings offer new molecular insights that could inform precision diagnosis and personalized treatment strategies for diabetic kidney disease.

Diabetic Nephropathies↗

Metabolic reprogramming and taxonomic drivers in bacterial vaginosis: A large-scale metagenomic meta-analysis.

OBJECTIVE: Bacterial vaginosis (BV) represents a profound ecological shift from a Lactobacillus-dominated microbiota to a diverse polymicrobial biofilm associated with adverse outcomes. While taxonomic signatures are well-documented, the functional mechanisms driving this transition remain obscured. This study elucidates the genomic potential for metabolic reprogramming and the putative "functional handover" underpinning the stability of the dysbiotic state. METHODS: A computational meta-analysis of 3557 vaginal microbiomes from diverse global cohorts was performed using the standardized MGnify pipeline. A high-resolution subset of 187 whole-genome shotgun (WGS) metagenomes was stratified to compare functional potential across demographic groups. Taxon-function interaction networks were constructed, utilizing a dual-filter statistical approach (p&#x202f;<&#x202f;0.05 and effect size ranking), to map the shift from homeostatic maintenance to dysbiotic metabolic potential. RESULTS: BV was characterized by a fundamental shift from "maintenance" pathways to high-turnover "growth-oriented" genomic repertoires. While ABC transporter-like domains were present in healthy communities, dysbiosis was marked by a quantitative expansion and diversification of these systems alongside P-loop NTPases. Network analysis revealed a putative "functional handover": while Gardnerella serves as the adherent structural scaffold, the metabolic burden appears to be associated with secondary anaerobes, specifically BVAB1 and Sneathia, which exhibit strong genomic correlations with nutrient transport and stress response pathways. Crucially, microbiomes from women of African ancestry (Black cohort) exhibited a distinct functional profile with genomic signatures consistent with functions previously associated with resistome expansion (e.g., tetracycline/macrolide resistance), contrasting with Asian cohorts. CONCLUSION: BV is a state of metabolic reprogramming where genomic functional dominance is transferred from Lactobacillus to a cooperative network of anaerobic opportunists. Identifying BVAB1 and Sneathia as candidate metabolic engines, supported by a Gardnerella scaffold, challenges current therapeutic paradigms and highlights the potential for precision medicine targeting specific functional drivers and resistome profiles across diverse populations.

Humans↗

Discovery of diverse anellovirus sequences in Thai human sequencing data.

UNLABELLED: Anelloviruses are part of the normal human viral flora. Although their diversity in humans has been investigated in many countries, and despite their initial detection in Thailand in 1999, knowledge of Thai anelloviruses remains very limited. This study analyzed 1,175 whole-genome sequencing data sets from Thai individuals to mine for potential anellovirus sequences. Our analyses detected anellovirus sequences in 149 data sets (12.68%), uncovering 434 partial anellovirus sequences and 77 complete genome sequences, characterized by the presence of terminal redundancy, complete orf1, and the conserved untranslated region upstream of the orf1 gene. Sequence analyses indicated that these viruses belong to seven genera, including Alphatorquevirus, Betatorquevirus, Gammatorquevirus, Hetorquevirus, Lamedtorquevirus, Samektorquevirus, and Yodtorquevirus. Notably, Hetorquevirus, Lamedtorquevirus, Samektorquevirus, and Yodtorquevirus had not previously been reported in Thailand. Phylogenetic analysis of ORF1 protein sequences showed that Thai anelloviruses form multiple phylogenetic clusters with non-Thai anelloviruses, indicating frequent cross-country transmission and multiple origins of the virus in Thailand. Furthermore, sequence similarity network analysis identified 33 potentially novel anellovirus species in our data set. Our findings greatly expand the knowledge of anellovirus diversity in Thailand and demonstrate the potential of human whole-genome sequencing data as a valuable resource for viral discovery. Lastly, we highlight and discuss some challenges with the use of the current pairwise sequence similarity-based classification scheme, in particular, how gaps can influence similarity calculation and potentially lead to inconsistencies with a phylogenetic-based classification scheme. IMPORTANCE: Anelloviruses are widespread in humans, yet their diversity remains poorly characterized in many regions, including Thailand. Here, we demonstrate that human sequencing data sets, originally generated without the intention for virome research, can be effectively mined for anellovirus sequences, including complete genomes. Our findings reveal a substantial number of previously unreported anelloviruses in Thailand, significantly expanding the known diversity of the virus. We also highlight potential limitations of the current anellovirus species classification scheme, which is based on pairwise orf1 sequence similarity analysis with a hard threshold cutoff at 69%. Our results reveal that the current scheme can sometimes yield taxonomic groupings that are inconsistent with phylogenetic relationships, particularly when significant alignment gaps are present. Overall, our results show that existing human sequencing data can be effectively repurposed for virus discovery research and suggest the need for more robust and phylogenetically informed classification frameworks as viral sequence databases continue to expand.

Humans↗