Search PubMedSearch

SEARCH · Search PubMed

Results for “Sequencing Resource”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Gencube: centralized retrieval and integration of multi-omics resources from leading databases.

MOTIVATION: The volume of multi-omics data for diverse species is growing at an unprecedented rate, with new genome assemblies, related annotations, and high-throughput sequencing resources being submitted daily to various genomic data repositories. In response to this data influx, both existing and new databases are establishing optimized hierarchical structures to manage the vast amount of information. However, the lack of accessible command-line tools, combined with the functional limitations and unintuitive design of existing options, presents significant challenges for researchers. This gap underscores a critical need for a tool that enables streamlined retrieval and integration of omics data across these diverse repositories. RESULTS: We have developed Gencube, a command-line tool that enables centralized retrieval and integration of a comprehensive set of six different data types-genome assemblies, gene sets, annotations, sequences, comparative genomic data, and NGS-based omics resources-from various leading databases. AVAILABILITY AND IMPLEMENTATION: Gencube is a free and open-source tool, with its code available on GitHub: https://github.com/snu-cdrc/gencube and also archived on Zenodo: https://doi.org/10.5281/zenodo.14607649.

Databases, Genetic

Predicting Weight Loss After Vertical Sleeve Gastrectomy Using a Whole-genome Sequencing-derived Polygenic Risk Score in the All of Us Cohort.

OBJECTIVE: To create a genome-wide polygenic risk score (PRS) to improve prediction of a 12-month percentage weight loss (WL) after vertical sleeve gastrectomy (VSG). BACKGROUND: Variability in post-VSG WL is not well explained by clinical factors. The All of Us program provides access to a 414,830 short-read whole-genome sequencing resource, enabling unbiased discovery of genetic predictors after VSG. METHODS: VSG counts, demographic, anthropomorphic and vital sign information were obtained from the linked electronic health record. The discovery cohort (DC) included participants from version 7 carried into version 8 while the validation cohort (VC) included those newly added to v8. We defined good responders and nonresponders as having WL&#xb1;1SD from the mean. Following quality filtering, we applied a 2-stage penalized-regression, followed by elastic-net logistic regression, to identify 1583 stable variants and derive &#x3b2;-weights. We then tested this PRS on the DC into a prediction model. RESULTS: We identified 395 participants in the DC and 336 participants in the VC, respectively. Of these, VSG, 44 were classified as good responders (&#x2265;37% WL) and 55 as nonresponders (&#x2264;19% WL). In the VC, 55 were classified as good responders and 48 as nonresponders. Adding the PRS to models to clinical predictors increased the area under the curve following logistic regression by 0.03; P <4.3 &#xd7; 10 -14 , random forest by 0.03; P <9.1 &#xd7; 10 -7 , decision tree by 0.05; P = 1.2 &#xd7; 10 -3 , and gradient boosting by 0.08; P <8.3 &#xd7; 10 -10 . CONCLUSIONS: Use of short-read whole-genome sequencing from All of Us (AoU) can be effectively used to generate PRS to enhance predictive WL accuracy. This work has implications for outcomes of both bariatric surgery and other surgical procedures.

Humans

Genomics-informed approach identifies which cell types regulate the metabolome.

MOTIVATION: Metabolism occurs in a cell type-specific manner, but which cells regulate metabolite levels remains unclear. RESULTS: Here, we integrate some of the largest metabolite quantitative trait loci datasets, TOPMed and UK Biobank, with one of the most extensive single-cell RNA sequencing resources, Tabula Sapiens. This integration allows us to identify cell types that regulate metabolites body-wide. We find hepatocytes are the primary regulatory cell type for most metabolites, associating with 385/410 (94%) metabolites for whom an association is found. Additionally, our multi-gene approach reveals more metabolite associations with beta cells compared to those identified using a single-gene approach. For example, we identify novel metabolite-cell type associations, such as the association between phenylpropanoic acid and beta cells, this metabolite that was previously thought to be regulated by the microbiome. AVAILABILITY: Code used in this work is available via Github at https://github.com/haimkru/Metabolite-Cell-Type-Associations.

Metabolome

A note on a generalized single step theory for any number of hierarchical genomic matrices.

BACKGROUND: The Single Step algorithm allows combining information from genotyped and un-genotyped individuals, provided they are connected by a pedigree. However, current single step theory is limited to a single list of markers. RESULTS: We present a generalized single step (GSS) method that can accommodate any number of hierarchical molecular datasets (e.g. sequence, high and low density arrays) and pedigree, avoiding imputation. We prove that a similar efficient inversion algorithm exists. The method is recursive, starting with the highest marker density scenario. We illustrate the method with simulation and show that GSS can increase predictive accuracy compared to standard single step. R code is provided so that custom scenarios can be easily compared, either with simulated or real data. CONCLUSION: The method developed generalizes extant single step theory to any number of hierarchical molecular relationship matrices, broadening the scenarios where single step can be applied. A topic of particular interest can be ecology field data or human populations where pedigree is not available, but where samples sequenced and genotyped at different densities can exist. GSS can also be a useful tool to optimize allocation of genotyping and / or sequencing resources.

Algorithms

IMAGE cDNA clones, UniGene clustering, and ACeDB: an integrated resource for expressed sequence information.

In this study we describe a new information resource that provides integrated access to information on IMAGE (integrated molecular analysis of genomes and their expression) cDNA library clones and derived expressed sequence tags (ESTs). We have developed an automated procedure that collates data from various public sources into a single ACeDB database. This database is a valuable tool for electronic cloning experiments and gene expression studies. It allows researchers to find information about cDNA libraries, plate addresses, insert sizes, and sequence data for IMAGE clones, the assignment of ESTs to UniGene clusters, and the chromosomal location of those genes in an efficient, graphically oriented manner.

Cloning, Molecular

Sequencing and health data resource of children of African ancestry.

PURPOSE: Individuals who self-report as Black or African American are historically underrepresented in genome-wide studies of disease risk, a disparity particularly evident in pediatric disease research. To address this gap, Cincinnati Children's Hospital Medical Center (CCHMC) established a biorepository and developed a comprehensive DNA sequencing resource including 15,684 individuals who self-identified as African American or Black and received care at CCHMC. METHODS: Participants were enrolled through the CCHMC Discover Together Biobank and sequenced. Admixture analyses confirmed the genetic ancestry of the cohort, which was then linked to electronic medical records. RESULTS: Genome-wide genotypes from common variants accompanied by medical record-sourced data are available through the Genomic Information Commons. This data set performs well in genetic studies. Specifically, we replicated known associations in sickle-cell disorder (HBB, HGNC:4827, P = 4.05 &#xd7; 10-148), anxiety (PLAAT3, HGNC:17825, P = 6.93 &#xd7; 10-9), and asthma (PCDH15, HGNC:14674, P = 5.6 &#xd7; 10-10), while also identifying novel loci associated with anxiety, asthma, and asthma severity. CONCLUSION: We present the acquisition and quality of genetic and disease-associated data and present an analytical framework for using this resource. In partnership with a community advisory council, we have codeveloped a valuable framework for data use and future research.

Adolescent

Uncovering new lineages in the Sunda pangolin (Manis javanica) with museum mitogenomics.

Accurately identifying evolutionarily significant units (ESUs) is crucial for conservation planning, especially for species like pangolins threatened by overhunting and habitat loss. ESUs help categorize different pangolin populations, aiding in understanding their genetic diversity and distribution, which is vital for targeted conservation efforts. This research generated mitochondrial genomes from historical museum specimens of Sunda pangolins (Manis javanica) from underrepresented locations, uncovering a new evolutionary lineage from the Mentawai Islands that diverged from Indochina and west Sundaland populations around 760 000 years ago. This population thereby represents a divergent ESU with a small distribution, important for conservation planning. The novel sequences provide resources for forensic labs tracing the origin of confiscated scales and shed light into the potential distribution of the 'mysterious pangolin'. Additionally, this research confirmed the presence of the two major M. javanica lineages in Java and extended the known distribution of the eastern clade to Bali and East Kalimantan. Our findings potentially suggest a recent bottleneck and postglacial expansion of pangolins across Indochina and west Sundaland. Further investigation with genomic and morphological evidence, contact area sampling and type sequencing will be required to evaluate the taxonomic status of different M. javanica lineages and M. culionensis.

Genomics

The Biobank Rare Variant consortium powers the discovery of rare genetic associations through global collaboration.

Rare coding variants can have large effects on disease risk and provide direct routes from human genetics to disease mechanisms and therapeutic targets, but their discovery is constrained by sample size, particularly for low-prevalence diseases. Here we establish the Biobank Rare Variant Analysis (BRaVa) consortium, a global rare variant association resource that integrates sequencing and linked health-record data from ten biobanks and cohorts comprising over 1.2 million individuals across diverse ancestries. We performed gene-based meta-analyses of rare coding variation across 33 clinical endpoints and 11 quantitative traits. Aggregating evidence across biobanks and ancestries identified 514 gene-trait associations, including 31 not previously reported in prior studies or curated association resources following systematic literature review. Notably, 36.1% of gene-level associations were undetectable in any individual biobank, and 91 emerged only through cross-ancestry meta-analysis, demonstrating that federated integration enables discovery beyond the reach of single cohorts. Similar gains were observed at the variant level, where 25.0% of phenotype-locus associations were detectable only through meta-analysis. Effect size estimates were correlated across ancestries with concordant directions of effect, supporting the generalizability of rare variant associations. The identified signals implicate pathways involved in transcriptional and epigenetic regulation, metabolism, vascular and epithelial biology, and immune function, highlighting rare coding variation as an engine for biological discovery across medical record phenotypes. For example, damaging variation in ANKRD12 implicates inflammatory transcriptional dysregulation in asthma and chronic obstructive pulmonary disease, and ultra-rare predicted loss-of-function variants in NAA15 link protein acetylation processes to type 2 diabetes risk. BRaVa establishes a scalable framework and freely available community resource for rare variant meta-analysis across global biobanks. Public release of gene- and variant-level association summary statistics provides a reference map of rare coding variant associations to support disease gene discovery, biological interpretation, and therapeutic target prioritization as sequencing-linked health-record resources continue to expand.

Journal Article

Mitochondrial genome-derived microsatellites reveal genetic diversity and population structure in Callery pear populations.

Callery pear (Pyrus calleryana Decne.; PC) possesses many desirable characteristics valued in managed landscapes. This has driven the release of numerous cultivars, including both hybrids and selections derived from native populations. The extensive planting of PC cultivars in managed areas has contributed to the widespread occurrence of invasive individuals across a broad range of habitats in the eastern United States (US). Self-incompatibility, tolerance to various environmental conditions, pathogen and pest resistance, intraspecific hybridization among the cultivars, possible interspecific hybridization with other Pyrus species, and seed dispersal by various vertebrates have contributed to the spread and persistence of PC across diverse environments. Because effective and environmentally appropriate management options remain limited, improved understanding of PC genetics may help inform management strategies. Previous studies have characterized PC diversity using nuclear genomic short sequence repeats (gSSRs), however, neither a mitochondrial genome resource nor mitochondrial short sequence repeats (mtSSRs) have been developed for this purpose. Here, we assembled a mitochondrial genome of 485,892 bp and used five mtSSRs to characterize mitochondrial&#xa0;diversity and population structure among accessions from the species' native range in Asia (n&#x2009;=&#x2009;72), southeastern US escapees (SNesc; n&#x2009;=&#x2009;90), Tennessee escapees (TNesc; n&#x2009;=&#x2009;90), and US-released commercial cultivars (UScult; n&#x2009;=&#x2009;69 representing 14 unique cultivars). We found a high genetic diversity (He&#x2009;=&#x2009;0.728) and evidence of genetic structure in PC. In distance-based and multivariate analyses, UScult occupied an intermediate position between the Asian populations and the US escapees. The observed mitochondrial diversity among samples assigned to PC cultivars is consistent with a complex genetic landscape and may reflect distinct maternal lineages, cultivar-labeling or record-keeping discrepancies, and/or technical variation. This study underscores the need for broader genomic investigations using authenticated cultivar reference material and high-resolution nuclear markers to resolve cultivar ancestry, validate true-to-name identity, and inform species management.

Genetic Variation

Molecular breeding of tomato: Advances and challenges.

The modern cultivated tomato (Solanum lycopersicum) was domesticated from Solanum pimpinellifolium native to the Andes Mountains of South America through a "two-step domestication" process. It was introduced to Europe in the 16th century and later widely cultivated worldwide. Since the late 19th century, breeders, guided by modern genetics, breeding science, and statistical theory, have improved tomatoes into an important fruit and vegetable crop that serves both fresh consumption and processing needs, satisfying diverse consumer demands. Over the past three decades, advancements in modern crop molecular breeding technologies, represented by molecular marker technology, genome sequencing, and genome editing, have significantly transformed tomato breeding paradigms. This article reviews the research progress in the field of tomato molecular breeding, encompassing genome sequencing of germplasm resources, the identification of functional genes for agronomic traits, and the development of key molecular breeding technologies. Based on these advancements, we also discuss the major challenges and perspectives in this field.

Solanum lycopersicum

Genomic characteristics and tracing analysis of an acute gastroenteritis outbreak associated with rotavirus C in a boarding high school.

BACKGROUND: Rotaviruses are major pathogens of childhood acute gastroenteritis, dominated by rotavirus A (RVA). Outbreaks caused by human rotavirus C (RVC) are rarely reported, and relevant genomic data remain scarce. This genomic investigation of an RVC outbreak improves our understanding of viral diversity and transmission dynamics. METHODS: We performed epidemiological surveys, nucleic acid testing and whole-genome sequencing (WGS) on specimens from a 2025 RVC-associated gastroenteritis outbreak at a Chinese boarding high school. Sequence alignment, phylogenetic and molecular tracing analyses were conducted to explore RVC evolution via point mutation, segment reassortment and genomic recombination. RESULTS: This typical point-source campus outbreak was linked to an indoor student gathering matching the incubation period of RVC. Thirteen RVC FX strains were recovered from 11 rectal swabs and two vomitus samples. Their viral protein (VP) 4 and VP7 sequences shared high homology with Russian reference strains, carrying distinct amino acid variations. No segment reassortment or recombination was detected in VP4/VP7 genes. CONCLUSIONS: Dense, closed campus settings facilitate RVC clustered transmission. Limitations included absent screening of asymptomatic canteen staff. Rapid nucleic acid testing enabled timely pathogen identification for outbreak control. Greater attention should be paid to the public health risk of RVC. These whole-genome sequencing data enrich resources for studying RVC evolution and vaccine development.

Acute gastroenteritis outbreak

Pan-genome based on chromosome sequences of wild and cultivated Agaricus bisporus.

Agaricus bisporus, one of the most widely cultivated mushrooms around the world, plays an important role in economy and agriculture. In this study, by employing long-reads generated by PacBio and Nanopore sequencing, we assembled six novel high-quality genomes (of which three are telomere-to-telomere assemblies) with sizes 29.6 ~ 30.8&#x2009;Mb and N50 lengths of 2.5 ~ 2.6&#x2009;Mb. Combined with public genome data of nine strains, we successfully established a pan-genome of A. bisporus, comprising a total of 14,626 clusters of protein coding genes, of which 50.70%, 7.45%, 24.74% and 17.01% are defined as core, soft core, dispensable, and private clusters, respectively. A total of 5,646 non- redundant structural variants (SVs) were identified among wild and cultivated strains and the genes associated with SV were mapped. This work provides valuable whole-genome sequences and genomic resources across wild and cultivated strains of the most widely cultivated mushroom species for functional analyses of genomes.

Agaricus

BAV-LLPS: a database of bacterial, archaea, and virus liquid-liquid phase separation proteins.

MOTIVATION: Liquid-liquid phase separation (LLPS) is a key process underlying the formation of biomolecular condensates, such as membrane-less organelles, that compartmentalize biochemical processes inside the cells. While LLPS has been extensively studied in eukaryotes, its role in bacteria, archaea, and viruses remains far less characterized. Recent studies in bacteria have revealed that LLPS-driven condensates play critical roles in RNA processing, stress response, and pathogenicity. Similarly, many viruses exploit LLPS to facilitate crucial steps in their infection cycles, including viral entry, genome replication, assembly, and host immune evasion. RESULTS: In this work, we introduce a hand-curated database of LLPS proteins from bacteria, archaea, and viruses (BAV-LLPS Database). This resource, extended through sequence similarity searches, comprises over 5000 proteins and integrates diverse data including biological annotations, sequence features, predicted disordered regions, LLPS per site probability, and AlphaFold2-based structural models. Additionally, our web server enables users to explore both the curated and homologous derived datasets, providing a platform to uncover evolutionary relationships and intrinsic and differential properties of LLPS proteins across various taxonomic groups. This work seeks to deepen our understanding of LLPS mechanisms beyond eukaryotic organisms, emphasizing their significance across diverse life forms. It also aims to foster the development of specialized predictive tools that will facilitate the exploration and characterization of LLPS processes in a wide array of living organisms, thereby contributing to advancements in both fundamental biological research and applied biomedical sciences. AVAILABILITY AND IMPLEMENTATION: BAV-LLPS DB is freely accessible at https://bav-llps-db.bioinformatica.org/. The data can be retrieved from the website. The source code of the database can be downloaded from https://bav-llps-db.bioinformatica.org/download.

Databases, Protein

A trainable language model with potential to modulate translation rates in non-model organisms by generating upstream untranslated region sequence libraries.

Tuning protein expression in non-model organisms is often constrained by the lack of validated genetic parts and predictive design tools. Translational tuning through the modulation of upstream untranslated regions (5'-UTRs) offers a potentially organism-agnostic route, but existing methods typically rely on mechanistic assumptions, prior knowledge that may not be available in non-model contexts, or the screening of sequence libraries. Here, we present a simple generative approach for creating synthetic 5'-UTR libraries based solely on the genomic sequence statistics of any desired organism. The method uses a sliding-window n-gram language model applied to native 5'-UTR sequences to produce novel sequences that preserve organism-specific base distributions and motifs without hard-coding specific motifs or mechanistic rules into inflexible statistical templates. We have applied this approach to the model bacterium Escherichia coli and the non-model probiotic Limosilactobacillus reuteri. Libraries of approximately 1,000 sequences were generated for each organism, from which about 100 unique sequences were experimentally tested for translation of a fluorescent reporter protein. In both organisms, the synthetic libraries yielded a broad range of translation levels from this relatively small number of tested variants. Sequences derived from an organism's own genomic statistics provided a more uniformly distributed range of translation rates in that organism than sequences derived from the other species. Correlations of individual sequence performance across the two species were weak, and thermodynamic predictions of ribosome binding strength showed very little predictive power, especially in the non-model L. reuteri. The results demonstrate that simple statistical language model approaches applied to genomic data can generate functional translational regulatory sequence libraries without detailed mechanistic knowledge or explicit reference to consensus motifs. The approach requires minimal computational resources, avoids reproducing native sequences, and can be readily applied to any organism with a sequenced genome. This strategy may lower technical barriers to expression tuning in non-model organisms.

5' Untranslated Regions

Mining thermophile photosynthesis genes: a synthetic operon expressing Chloroflexota species reaction center genes in Rhodobacter sphaeroides.

Photosynthesis is the foundation of the vast majority of life systems, and therefore the most important bioenergetic process on earth, and the greatest diversity in photosynthetic systems are found in microorganisms. However, understanding of the biophysical and biochemical processes that transduce light to chemical energy has derived from the relatively small subset of proteins from microbes that are amenable to cultivation, in contrast to the huge number of microbial DNA sequences encoding proteins that catalyze the initial photochemical reactions that has been deposited in databases, such as from metagenomics. We describe the use of a Rhodobacter sphaeroides laboratory strain for expression of heterologous photosynthesis genes to demonstrate the feasibility of mining this resource, focusing on hot spring Chloroflexota gene sequences. Using a synthetic operon of genes, we produced a photochemically active complex of reaction center proteins in our biological system. We also present bioinformatic analyses of anoxygenic type II reaction center sequences from metagenomic samples collected from hot (42-90&#xb0; C) springs available through the JGI IMG database, to generate a resource of diverse sequences that potentially are adapted to photosynthesis at such temperatures. These data provide a view into the natural diversity of anoxygenic photosynthesis, through a lens focused on high-temperature environments. The approach we took to express such genes can be applied for potential biotechnology purposes as well as for studies of fundamental catalytic properties of these heretofore inaccessible protein complexes.

Chloroflexota

Expanding vaginal microbiome pangenomes via a custom MIDAS database reveals Lactobacillus crispatus accessory genes associated with cervical dysplasia.

The vaginal microbiome plays a central role in reproductive health. Vaginal microbiome dysbiosis is associated with many adverse reproductive health outcomes, but most studies have focused on associations at the species level. The potential contribution of intraspecies microbial variation, especially gene content differences across bacterial strains, remains underexplored in reproductive health contexts. The Metagenomic Intra-Species Diversity Analysis (MIDAS) framework enables such analyses, but depends on comprehensive reference databases. We constructed a MIDAS-compatible pangenome database from over 18,000 genomes in the Vaginal Microbiome Genome Collection (VMGC). Compared to the Genome Taxonomy Database (GTDB)-derived reference, the VMGC-derived database expanded the pangenomes of prevalent vaginal species, better capturing vaginal-specific intraspecies diversity. Applying this database to vaginal samples from a cervical dysplasia cohort, we identified 13 Lactobacillus crispatus accessory genes significantly associated with cervical dysplasia, including a HicAB toxin-antitoxin system, three transcriptional regulators, and three phage-derived genes. These findings highlight the utility of body site-specific reference resources and shotgun metagenomic sequencing for uncovering intraspecies microbial variation relevant to reproductive health.IMPORTANCEThe vaginal microbiome plays a critical role in reproductive health, and different bacteria from the same species can carry different genes that influence how the strains interact with the host and other microbes. These strain-level differences are often overlooked when microbiomes are analyzed only at the species level. Existing genomic reference databases are heavily biased toward gut and environmental bacteria, leaving the genetic diversity of vaginal microbes understudied. We built a specialized reference database from over 18,000 vaginal bacterial genomes that better reflects this diversity. We then applied this resource to quantify gene-level variation in vaginal samples from a cervical dysplasia cohort. Focusing on Lactobacillus crispatus, a prevalent and often beneficial vaginal species, we identified 13 genes that were more common in women with cervical dysplasia than in controls. This work demonstrates that body site-specific genomic resources are essential for uncovering strain-level bacterial differences relevant to reproductive health.

Lactobacillus crispatus

A comparison of software for analysis of rare and common short tandem repeat (STR) variation using human genome sequences from clinical and population-based samples.

Short tandem repeat (STR) variation is an often overlooked source of variation between genomes. STRs comprise about 3% of the human genome and are highly polymorphic. Some cause Mendelian disease, and others affect gene expression. Their contribution to common disease is not well-understood, but recent software tools designed to genotype STRs using short read sequencing data will help address this. Here, we compare software that genotypes common STRs and rarer STR expansions genome-wide, with the aim of applying them to population-scale genomes. By using the Genome-In-A-Bottle (GIAB) consortium and 1000 Genomes Project short-read sequencing data, we compare performance in terms of sequence length, depth, computing resources needed, genotyping accuracy and number of STRs genotyped. To ensure broad applicability of our findings, we also measure genotyping performance against a set of genomes from clinical samples with known STR expansions, and a set of STRs commonly used for forensic identification. We find that HipSTR, ExpansionHunter and GangSTR perform well in genotyping common STRs, including the CODIS 13 core STRs used for forensic analysis. GangSTR and ExpansionHunter outperform HipSTR for genotyping call rate and memory usage. ExpansionHunter denovo (EHdn), STRling and GangSTR outperformed STRetch for detecting expanded STRs, and EHdn and STRling used considerably less processor time compared to GangSTR. Analysis on shared genomic sequence data provided by the GIAB consortium allows future performance comparisons of new software approaches on a common set of data, facilitating comparisons and allowing researchers to choose the best software that fulfils their needs.

Humans