Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

The genome sequence of the Yellow-barred Brindle, Acasis viretata (Hübner, 1799) (Lepidoptera: Geometridae).

We present a genome assembly from an individual female Acasis viretata (Yellow-barred Brindle; Arthropoda; Insecta; Lepidoptera; Geometridae). The genome sequence has a total length of 297.68 megabases. Most of the assembly (99.98%) is scaffolded into 17 chromosomal pseudomolecules, including the W and Z sex chromosomes. The mitochondrial genome has also been assembled, with a length of 16.01 kilobases. This assembly was generated as part of the Darwin Tree of Life project, which produces reference genomes for eukaryotic species found in Britain and Ireland.

Acasis viretata; Yellow-barred Brindle; genome seq↗

The genome sequence of an ichneumonid wasp, Netelia melanura (Thomson, 1888) (Hymenoptera: Ichneumonidae).

We present a genome assembly from an individual male Netelia melanura (ichneumonid wasp; Arthropoda; Insecta; Hymenoptera; Ichneumonidae). The genome sequence has a total length of 253.87 megabases. Most of the assembly (94.81%) is scaffolded into 7 chromosomal pseudomolecules. The mitochondrial genome has also been assembled, with a length of 28.04 kilobases. This assembly was generated as part of the Darwin Tree of Life project, which produces reference genomes for eukaryotic species found in Britain and Ireland.

Netelia melanura; ichneumonid wasp; genome sequenc↗

The genome sequence of a window fly, Scenopinus jerei Pohjoismäki & Haarto, 2021 (Diptera: Scenopinidae).

We present a genome assembly from an individual male Scenopinus jerei (window fly; Arthropoda; Insecta; Diptera; Scenopinidae). The assembly contains two haplotypes with total lengths of 345.25 megabases and 232.44 megabases. Most of haplotype 1 (94.85%) is scaffolded into 5 chromosomal pseudomolecules, including the X and Y sex chromosomes. Haplotype 2 was assembled to scaffold level. The mitochondrial genome has also been assembled, with a length of 16.52 kilobases. This assembly was generated as part of the Darwin Tree of Life project, which produces reference genomes for eukaryotic species found in Britain and Ireland.

Scenopinus jerei; window fly; genome sequence; chr↗

The genome sequence of the Orange Underwing, Archiearis parthenias (Linnaeus, 1761) (Lepidoptera: Geometridae).

We present a genome assembly from an individual female Archiearis parthenias (Orange Underwing; Arthropoda; Insecta; Lepidoptera; Geometridae). The assembly contains two haplotypes with total lengths of 528.46 megabases and 426.17 megabases. Most of haplotype 1 (99.61%) is scaffolded into 26 chromosomal pseudomolecules, including the W and Z sex chromosomes. Haplotype 2 was assembled to scaffold level. The mitochondrial genome has also been assembled, with a length of 17.24 kilobases. This assembly was generated as part of the Darwin Tree of Life project, which produces reference genomes for eukaryotic species found in Britain and Ireland.

Archiearis parthenias↗

The genome sequence of a blow fly, Melanomya nana (Meigen, 1826) (Diptera: Calliphoridae).

We present a genome assembly from an individual male Melanomya nana (blow fly; Arthropoda; Insecta; Diptera; Calliphoridae). The assembly contains two haplotypes with total lengths of 1 391.98 megabases and 1 424.34 megabases. Most of haplotype 1 (99.62%) is scaffolded into 5 chromosomal pseudomolecules. Haplotype 2 was assembled to scaffold level. The mitochondrial genome has also been assembled, with a length of 27.12 kilobases. This assembly was generated as part of the Darwin Tree of Life project, which produces reference genomes for eukaryotic species found in Britain and Ireland.

Diptera↗

Genome-Wide Identification of SSR and InDel Markers and Experimental Validation of SSR Markers for Distinguishing Cold-Tolerant and Cold-Sensitive Lily Cultivars.

In this study, whole-genome resequencing was performed on the cold-tolerant variety ND-6 and the cold-sensitive variety 'Sorbonne'. After evaluation, the Lilium davidii var. unicolor reference genome was selected to analyze SSR distribution characteristics. Whole-genome InDel identification and comparative analysis were conducted for the two varieties, yielding 34,812,909 and 24,497,857 InDels, respectively. Short InDels were predominant, with deletions slightly outnumbering insertions, mostly located in intergenic regions. Twenty pairs of SSR primers were screened and synthesized. Among them, 10 pairs amplified clearly, with a polymorphism rate of 82.6%, effectively distinguishing the two cultivars examined in this study. This study provides systematic data and a reliable marker resource for the analysis of lily genomic variation, laying a foundation for the identification of cold-tolerant germplasm; validation across additional cultivars and individuals will be required to extend their utility to broader germplasm.

cold resistant lilies↗

Tandem splice acceptor sites: Profiling their relevance to human disease.

PURPOSE: Interpretation of variation, particularly the creation or disruption of tandem splice acceptor sites (NAGNnAG variants), challenges genomic medicine practice. METHODS: We analyzed the creation and disruption of dinucleotide AG sites within ±30 bases of natural splice-acceptor sites in the GRCh37 human reference genome. These results were compared with variant data from the ClinVar and gnomAD databases, as well as with data from 779 National Institutes of Health Undiagnosed Diseases Program study participants. Using RNA sequencing, we assessed the splicing at NAGNnAG variants for 107 of the Undiagnosed Diseases Program participants and compared the empirical data with SpliceAI predictions. RESULTS: Creation or disruption of NAGNnAG sites within 30 bases of the natural splice acceptor are enriched in ClinVar compared with gnomAD; however, such variants in the 2 databases are rarely differentiated by SpliceAI scores. Empirical evaluation via RNA sequencing analysis supported novel acceptor site usage from -21 to +30; splice-altering variants did not predominate in a specific region or have SpliceAI scores invariantly, suggesting increased spliceogenicity. CONCLUSION: NAGNnAG variants within 30 bp of the natural splice acceptor have a high probability of clinical relevance and are poorly contextualized for clinical utility. Their interpretation benefits from empirical evaluation via RNA analysis.

Humans↗

Epidemiological and phylogenetic analysis of anthrax in Kazakhstan in 2024.

BACKGROUND: Anthrax remains an important zoonotic disease in Kazakhstan due to the persistence of environmental reservoirs and long-standing endemic foci. Despite ongoing surveillance, the epidemiological characteristics and genetic diversity of circulating Bacillus anthracis strains in the country remain incompletely understood. METHODS: A retrospective epidemiological and phylogenetic investigation of anthrax outbreaks reported in Kazakhstan during 2024 was conducted. Epidemiological data were collected for all laboratory-confirmed human cases and associated outbreak foci. Confirmation of infection was performed by PCR, and B. anthracis isolates were obtained from clinical, environmental and animal-associated samples. Whole-genome sequencing and core-genome single nucleotide polymorphism (cgSNP) analysis were used to characterize the genetic relationships among isolates and to determine their phylogenetic placement. RESULTS: Nine anthrax outbreaks were identified across four regions of Kazakhstan (Almaty, Zhambyl, Atyrau, and West Kazakhstan), resulting in 20 confirmed human cases. All patients were male, with the highest proportion occurring among individuals aged 36-55 years (45%). The mean patient age was 43.9 years (range: 16-64 years). Most infections were associated with slaughtering infected livestock (65%), followed by handling contaminated meat (15%). PCR confirmed infection in all 20 human cases. Culture yielded 17 human-derived B. anthracis isolates from 14 patients and 17 environmental/animal-derived isolates, resulting in 34 isolates in total. Of these, 22 representative isolates underwent whole-genome sequencing. Phylogenetic analysis revealed the circulation of two major lineages. Isolates from Atyrau and West Kazakhstan clustered within the Trans-Eurasian (TEA/STI) lineage. Atyrau isolates formed a tight cluster differing by only 21-32 cgSNPs, consistent with a shared epidemiolocal source, whereas the West Kazakhstan isolate was highly divergent. Zhambyl and Almaty region belonged to the A.Br.Ames lineage but diverged into two distinct sublineages. Zhambyl region isolates demonstrated minimal divergence from the global reference genome Ames Ancestor, differing by only 16-31 SNPs. Almaty region isolates formed an endemic subclone, separated from the reference group by approximately 114 SNPs. Comparison with the Ames Ancestor and Sterne reference strains demonstrated substantial genetic divergence. CONCLUSION: Anthrax outbreaks in Kazakhstan during 2024 were primarily associated with livestock exposure and occurred within established endemic regions. Whole-genome sequencing revealed the coexistence of distinct TEA and Ames lineages, including evidence of persistent local transmission and long-term evolutionary stability of endemic B. anthracis populations. These findings enhance understanding of anthrax epidemiology in Central Asia and support the integration of genomic surveillance into national outbreak investigation programs.

Anthrax↗

Efficient high-throughput resequencing of genomic DNA.

Targeted resequencing of genomic DNA from organisms such as humans is an important tool enabling experimental access to variation within the species and between similar species. Taking full advantage of the reference genome sequences in designing robust, specific PCR assays and using stringent conditions, resequencing can be done efficiently without purification of the PCR product. By using a 10-fold greater amount of one primer when setting up the PCR initially in a new version of asymmetric PCR, one simply adds the rest of the sequencing reagents at the end of PCR and allows the sequencing reaction to proceed, with the excess PCR primer serving as the sequencing primer. We demonstrated that this streamlined protocol can be used with PCR products up to 1300 bp and had up to a 97% success rate in high-throughput analysis of allele frequencies for >30,000 single-nucleotide polymorphisms (SNPs). SNP primers and characterization results are provided at http://snp.wustl.edu.

Alleles↗

Integrative genomic, transcriptional, and proteomic diversity in natural isolates of the human pathogen Burkholderia pseudomallei.

Natural isolates of pathogenic bacteria can exhibit a broad range of phenotypic traits. To investigate the molecular mechanisms contributing to such phenotypic variability, we compared the genomes, transcriptomes, and proteomes of two natural isolates of the gram-negative bacterium Burkholderia pseudomallei, the causative agent of the human disease melioidosis. Significant intrinsic genomic, transcriptional, and proteomic variations were observed between the two strains involving genes of diverse functions. We identified 16 strain-specific regions in the B. pseudomallei K96243 reference genome, and for eight regions their differential presence could be ascribed to either DNA acquisition or loss. A remarkable 43% of the transcriptional differences between the strains could be attributed to genes that were differentially present between K96243 and Bp15682, demonstrating the importance of lateral gene transfer or gene loss events in contributing to pathogen diversity at the gene expression level. Proteins expressed in a strain-specific manner were similarly correlated at the gene expression level, but up to 38% of the global proteomic variation between strains comprised proteins expressed in both strains but associated with strain-specific protein isoforms. Collectively, >65 hypothetical genes were transcriptionally or proteomically expressed, supporting their bona fide biological presence. Our results provide, for the first time, an integrated framework for classifying the repertoire of natural variations existing at distinct molecular levels for an important human pathogen.

Bacterial Proteins↗

Chromosome-level genome assembly of Sinocyclocheilus jii based on PacBio HiFi and Hi-C sequencing.

Sinocyclocheilus jii, a cavefish species endemic to China, belongs to the genus Sinocyclocheilus within the family Cyprinidae. Species within this genus exhibit significant morphological differentiation, making it not only the most species-rich genus within Cyprinidae in China but also the most diverse group of cavefishes worldwide. However, the limited availability of genomic resources has limited investigations into the genetic basis of trait variations, phylogenetic relationships, and adaptive evolution in this genus. In this study, we assembled a chromosome-level reference genome for S. jii by integrating PacBio HiFi long reads, Illumina short reads, and Hi-C sequencing data. Flow cytometry was used to estimate the genome size prior to assembly, providing a key step in technical validation. The final genome assembly spans 1.75 Gb with a contig N50 of 35.0 Mb. Using Hi-C sequencing data, the assembled scaffolds were successfully anchored to 50 chromosomes. The completeness of the chromosome-level assembly was estimated at 98.9% by BUSCO analysis. Genome annotation identified 855.5 Mb of repetitive sequences and predicted a total of 52,867 protein-coding genes, of which 51,932 genes were functionally annotated. This study presents a high-quality chromosome-level genome assembly and annotation of S. jii, providing a fundamental genomic resource for future phylogenetic and evolutionary studies.

Animals↗

National genomic projects in Asia and Africa: a review.

National genome projects (NGPs) are increasingly shaping precision medicine by improving representation of population-specific genetic diversity. This review compiles findings from NGPs across Asia and Africa, regions that remain underrepresented in global genomic databases despite their extensive demographic and genetic diversity. A total of 53 studies from 24 countries were identified to understand (1) the genomic approach utilized, (2) novel findings that have emerged, and (3) strategies for improving research in these regions. The NGPs implement population-based variome databases (20 NGPs), linear reference genome assemblies (8 NGPs), and graph-based pangenome assemblies (1 NGP). Novel variants ranged between 0.28% (China) and 19.6% (Iran), whereas rare variants accounted for up to 88.9% of the detected variants in the Chinese population. Each NGP documents its country's evolutionary and migration history, which impacts disease frequency and pharmacogenomic variants. Clinically, NGPs revealed strong population stratification in disease-associated and pharmacogenomic variants. For example, the GJB2 rs72474224 hearing-loss variant ranged from 13% in Vietnam and 12% in Hong Kong to 0.0894% in Turkey, while the VKORC1 rs9923231 pharmacogenomic variant reached 89.2% in Taiwan but was 20%-25% in European-related Russian subpopulations. These findings demonstrate that clinically relevant allele frequencies, pathogenicity assessments, and drug-response markers differ substantially across ancestries. This review highlights ongoing efforts and strategies to enhance the representativeness of genomic data through NGPs in Asia and Africa. We also suggest future directions for national projects, including integrating family-based studies, multi-omic data, and standardized pipelines to accelerate discovery and support the equitable implementation of precision medicine.

Humans↗

Parent-of-origin specific allelic expression in outbreeding Arabidopsis arenosa identifies antagonistic parental enrichment in protein degradation pathways.

In plants, the epigenetic phenomenon of parent-of-origin allele-specific expression occurs mainly in the triploid endosperm. Although well studied in inbreeding Arabidopsis thaliana, genomic imprinting has been less investigated in outcrossers. In order to investigate a wider role of parental-specific allelic expression, we have analyzed imprinting in whole seeds of the obligate outbreeder Arabidopsis arenosa. High-throughput analysis of imprinting in outbreeding species is hampered by the lack of reference genomes and available sequenced accessions. High degree of allelic variation in outbreeding species may also limit the analysis to loci with less variation. We developed a reference-independent pipeline to detect parental-specific reads. Using different accessions in reciprocal crosses, we detected more than 70 paternally biased imprinted genes and > 500 maternally biased genes. Paternally biased genes showed major enrichment for proteins with ubiquitin protein transferase and ligase activity. Maternally biased genes were enriched for protein pathways directly counteracting paternally enriched genes. Here, we demonstrate an alignment-free protocol to identify imprinted genes that may be successfully applied for imprinting studies in other highly heterozygous outcrossing species. Our results suggest a unique role of genomic imprinting affecting post-transcriptional gene regulation in outbreeding A. arenosa.

Arabidopsis arenosa↗

Alu Overexpression Leads to an Increased Double-Stranded RNA Signature in Dermatomyositis.

OBJECTIVE: Dermatomyositis is an autoimmune condition characterized by a high interferon signature of unknown etiology. Because coding sequences constitute <1.2% of our genomes, there is a need to explore the role of the noncoding genome in disease pathogenesis. Our genomes include roughly 1.2 million Alu elements occupying approximately 10% of the genome, which can form double-stranded (ds) RNA capable of triggering MDA5 leading to interferon production. METHODS: We aligned muscle biopsy RNA sequencing data to the telomere-to-telomere reference genome and quantified short interspersed elements including Alus. Because Alus have a propensity to form dsRNA and are the major targets of both adenosine deaminase RNA specific and MDA5, we quantified adenosine to inosine (A-to-I) RNA editing, which reflects dsRNA in vivo. RESULTS: Dermatomyositis muscle (n = 39) showed a global elevation in Alu expression (including inverted-repeat Alus with high potential to form dsRNA) as well as an increased expression of unique Alu elements (n = 557, q < 0.05) compared with healthy controls (n = 34), in a pattern not seen in other myositis types (n = 81). Most (75.3%) of these Alus originated from genomic regions outside genes. A cluster of the uniquely overexpressed Alus (n = 167) correlated with interferon-stimulated genes and markers of myositis activity. Additionally, we found a uniquely expanded Alu A-to-I editome in dermatomyositis, reflecting an increase in dsRNA. Edited Alus clustered on chromosome 19, which is known to have the highest concentration of dsRNA. CONCLUSION: We hypothesize that overexpressed Alus in dermatomyositis form endogenous dsRNA that exceeds the capacity of RNA editing enzymes and triggers dsRNA sensors leading to interferon production.

Humans↗

Apply innovative technologies to explore cancer genome.

PURPOSE OF REVIEW: Molecular genetic alterations characterize the development of human cancer. Recent advances in molecular genetic technology and the success of the human genome project have empowered investigators with new tools in dissecting the cancer genome for discovery of new cancer-associated genes. The purpose of this review is to highlight the emerging molecular genetic methodologies and summarize their principles, applications, and potential technical challenges. The critical issue in sample preparation and a strategy that combines different molecular techniques to facilitate the identification of novel cancer-associated genes will be discussed. RECENT FINDINGS: Digital karyotyping and array-based techniques including array comparative genomic hybridization and representational oligonucleotide microarray analysis have been recently developed to study the genomic landscape in human cancer. These innovations provide tools to quantitatively measure DNA copy number changes in cancer and to map those changes directly onto the human genome. Digital karyotyping is based on counting the sequence tags that are distributed in the human genome and thus, it provides a digital readout to precisely outline the amplified and deleted chromosomal regions. Array-based technologies, on the other hand, compare the content of cancer and reference genomes followed by localizing the amplified or deleted signals in chromosomal regions using an array hybridization technique. In addition, a high-throughput mutational analysis platform has been available for a large-scale mutational analysis by using an automated capillary sequencing device and sophisticated bioinformatic tools. A number of examples have demonstrated the promise of these new molecular genetic approaches in identifying several potential new oncogenes and tumor suppressors. SUMMARY: As compared with conventional cytogenetics methods, digital karyotyping, array comparative genomic hybridization, and representational oligonucleotide microarray analysis provide an unprecedented mapping resolution that allows a precise localization of the amplified and deleted chromosomal regions. These technologies can be combined with gene expression profiling and high-throughput mutational analysis to facilitate the search for new cancer-associated genes. It is expected that applying these new technologies will lead to discovery of a host of novel oncogenes and tumor suppressors, which will have a significant impact in our understanding of tumorigenesis and in the clinical management of cancer patients.

Cytogenetics↗

Chromosome-scale genome assembly and genomic prediction of essential oil compounds in Atractylodes lancea for genomics-assisted breeding.

Atractylodes lancea rhizomes are used as crude drugs. Essential oil compounds, including atractylodin, hinesol, &#x3b2;-eudesmol, and atractylon, are key determinants of crude drug quality. Conventional breeding of A. lancea is difficult because of its perennial growth. In this study, a chromosome-scale reference genome of A. lancea (4.79 Gb) was generated, and genome-wide association studies (GWAS) and genomic predictions of essential oil compounds were conducted to explore the potential for genome-assisted breeding. Genotyping of 480 lines using double-digest restriction-site-associated DNA-sequencing yielded 29,136 high-quality SNPs. All the compounds showed high genomic heritability (h2 = 0.758-0.915), indicating strong genetic control. Despite the high genomic heritability, GWAS detected only one weak association with atractylon and no significant loci for the three compounds. However, genomic prediction achieved moderate to high accuracy across multiple models, particularly the ridge regression, genomic best linear unbiased prediction, and Bayesian approaches. The prediction accuracy, measured as the Pearson correlation coefficient between the observed and predicted values, exceeded 0.6 for all four essential oil compounds. These results demonstrate the efficacy of genomic selection for improving essential oil compound levels in A. lancea and provide a foundation for genome-assisted breeding of medicinal plants with long breeding cycles.

Atractylodes lancea↗

Loss of heterozygosity and transcriptome analyses of a 1.2 Mb candidate ovarian cancer tumor suppressor locus region at 17q25.1-q25.2.

Loss of heterozygosity (LOH) analysis was performed in epithelial ovarian cancers (EOC) to further characterize a previously identified candidate tumor suppressor gene (TSG) region encompassing D17S801 at chromosomal region 17q25.1. LOH of at least one informative marker was observed for 100 (71%) of 140 malignant EOC samples in an analysis of 6 polymorphic markers (cen-D17S1839-D17S785-D17S1817-D17S801-D17S751-D17S722-tel). The combined LOH analysis revealed a 453 kilobase (Kb) minimal region of deletion (MRD) bounded by D17S1817 and D17S751. Human and mouse genome assemblies were used to resolve marker inconsistencies in the D17S1839-D17S722 interval and identify candidates. The region contains 32 known and strongly predicted genes, 9 of which overlap the MRD. The reference genomic sequences share nearly identical gene structures and the organization of the region is highly collinear. Although, the region does not show any large internal duplications, a 1.5 Kb inverted duplicated sequence of 87% nucleotide identity was observed in a 13 Kb region surrounding D17S801. Transcriptome analysis by Affymetrix GeneChip and reverse transcription (RT)-polymerase chain reaction (PCR) methods of 3 well characterized EOC cell lines and primary cultures of normal ovarian surface epithelial (NOSE) cells was performed with 32 candidates spanning D17S1839-D17S722 interval. RT-PCR analysis of 8 known or strongly predicted genes residing in the MRD in 10 EOC samples, that exhibited LOH of the MRD, identified FLJ22341 as a strong candidate TSG. The proximal repeat sequence of D17S801 occurs 8 Kb upstream of the putative promoter region of FLJ22341. RT-PCR analysis of the EOC samples and cell lines identified DKFZP434P0316 that maps proximal to the MRD, as a candidate. While Affymetrix technology was useful for initially eliminating less promising candidates, subsequent RT-PCR analysis of well-characterized EOC samples was essential to prioritize TSG candidates for further study.

Chromosome Mapping↗

Whole-Genome Sequencing of 54 Dengchuan Cattle (Bos taurus) from Southwest China.

Domestic cattle (Bos taurus) play a significant role in human society as they provide abundant food resources and contribute to the development of agriculture and traditional culture. Dengchuan cattle, a local breed from Yunnan, Southwest China, are known for their high-quality milk and are at risk of extinction due to crossbreeding. To preserve the superior genetic resources of Dengchuan cattle, this study conducted whole-genome sequencing of 54 Dengchuan cattle using blood DNA samples, generating approximately 3.56 TB of clean data with an average sequencing depth of 32.78X. The sequencing data were aligned to the bovine reference genome (ARS-UCD2.0), achieving an average alignment rate of 99.85%. A total of 9,950,420 SNPs and 2,476,207 indels were detected using variant calling workflow. These data were utilized to characterize genomic profile of this unique cattle breed. The data generated in this study can be incorporated into the global cattle genomic diversity database, providing valuable information for comparative studies on cattle.

Animals↗