Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Implications of human genome architecture for rearrangement-based disorders: the genomic basis of disease.

The term 'genomic disorder' refers to a disease that is caused by an alteration of the genome that results in complete loss, gain or disruption of the structural integrity of a dosage sensitive gene(s). In most of the common chromosome deletion/duplication syndromes, the rearranged genomic segments are flanked by large (usually >10 kb), highly homologous low copy repeat (LCR) structures that can act as recombination substrates. Recombination between non-allelic LCR copies, also known as non-allelic homologous recombination, can result in deletion or duplication of the intervening segment. Recent findings suggest that other chromosomal rearrangements, including reciprocal, Robertsonian and jumping translocations, inversions, isochromosomes and small marker chromosomes, may also involve susceptibility to rearrangement related to genome structure or architecture. In several cases, LCRs, AT-rich palindromes and pericentromeric repeats are located at such rearrangement breakpoints. Analysis of the products of recombination at the junctions of the rearrangements reveals both homologous recombination and non-homologous end joining as causative mechanisms. Thus, a more global concept of genomic disorders emerges in which susceptibility to rearrangements occurs due to underlying complex genomic architecture. Interestingly, this architecture plays a role not only in disease etiology, but also in primate genome evolution. In this review, we discuss recent advances regarding general mechanisms for the various rearrangements of our genome, and potential models for rearrangements with non-homologous breakpoint regions.

Biological Evolution↗

The genome of Salmonella enterica serovar gallinarum: distinct insertions/deletions and rare rearrangements.

Salmonella enterica serovar Gallinarum is a fowl-adapted pathogen, causing typhoid fever in chickens. It has the same antigenic formula (1,9,12:--:--) as S. enterica serovar Pullorum, which is also adapted to fowl but causes pullorum disease (diarrhea). The close relatedness but distinct pathogeneses make this pair of fowl pathogens good models for studies of bacterial genomic evolution and the way these organisms acquired pathogenicity. To locate and characterize the genomic differences between serovar Gallinarum and other salmonellae, we constructed a physical map of serovar Gallinarum strain SARB21 by using I-CeuI, XbaI, and AvrII with pulsed-field gel electrophoresis techniques. In the 4,740-kb genome, we located two insertions and six deletions relative to the genome of S. enterica serovar Typhimurium LT2, which we used as a reference Salmonella genome. Four of the genomic regions with reduced lengths corresponded to the four prophages in the genome of serovar Typhimurium LT2, and the others contained several smaller deletions relative to serovar Typhimurium LT2, including regions containing srfJ, std, and stj and gene clusters encoding a type I restriction system in serovar Typhimurium LT2. The map also revealed some rare rearrangements, including two inversions and several translocations. Further characterization of these insertions, deletions, and rearrangements will provide new insights into the molecular basis for the specific host-pathogen interactions and mechanisms of genomic evolution to create a new pathogen.

Bacterial Proteins↗

[Short introduction to gene imprinting].

Genomic imprinting refers to the genetic non-equivalence of mammalian paternal and maternal genomes. It leads to differential marking of gene alleles in parental gamets. This marking leads to differential expression of imprinted genes during embryonal development and in adult life. Usually one allele of an imprinted gene is active while the other is silent. Genomic imprinting renders mammals haploid for imprinted genes and thereby causes certain genetic diseases. These can be due to a variety of genetic events, such as loss or mutation of the active imprinted gene allele, disruption of imprinting in parental gamets so that both alleles are imprinted in the same way, or inheritance of a chromosome or part of chromosome bearing imprinted genes from a single parent. Several severe human genetic disorders are associated with imprinted genes, including Beckwith-Wiedemann, Prader-Willi and Angelman syndromes. In this review a glimpse of the complexity of genetic mechanisms, which lead to the differential transcription of imprinted genes, is presented.

English Abstract↗

Application of AFLP markers to genome mapping in poultry.

The amplified fragment length polymorphism (AFLP) technique has been used to enhance marker density in the East Lansing reference chicken genome map, using a backcross family derived from a Red Jungle Fowl by White Leghorn mating with White Leghorn as the recurrent parent. To date, 204 AFLP markers have been added, expanding overall map coverage by about 25%. To the limits of our resolution, AFLP markers are distributed relatively evenly across the EL reference map. AFLP are about 60% as frequent in a cross within White Leghorns (line 7(2) x 6(3)) in comparison to the more divergent reference map population. Based on apparent identity of size, about 40% of the 7(2) x 6(3) cross AFLP fragments were also polymorphic in the reference map cross. Primer pairs in which one primer contains 3' extensions of three selective nucleotides and the other has two selective nucleotides successfully generated AFLP from chicken DNA, but such pairs appeared to amplify only a subset of those fragments to which they have an exact sequence match. Three different restriction enzymes with 4 bp recognition sites (TaqI, HinP1I and MspI) were found to work well with EcoRI as the rarer of the two AFLP restriction enzymes used, with HinP1I being the most effective of the three. AFLP markers are likely to provide an economical method with which to enhance framework linkage maps of chicken and probably other avian genomes.

Animals↗

Shotgun metagenomic analysis of saliva microbiome suggests Mogibacterium as a factor associated with chronic bacterial osteomyelitis.

Osteomyelitis of the jaw is a severe inflammatory disorder that affects bones, and it is categorized into two main types: chronic bacterial and nonbacterial osteomyelitis. Although previous studies have investigated the association between these diseases and the oral microbiome, the specific taxa associated with each disease remain unknown. In this study, we conducted shotgun metagenome sequencing (≥10 Gb from ≥66,395,670 reads per sample) of bulk DNA extracted from saliva obtained from patients with chronic bacterial osteomyelitis (N = 5) and chronic nonbacterial osteomyelitis (N = 10). We then compared the taxonomic composition of the metagenome in terms of both taxonomic and sequence abundances with that of healthy controls (N = 5). Taxonomic profiling revealed a statistically significant increase in both the taxonomic and sequence abundance of Mogibacterium in cases of chronic bacterial osteomyelitis; however, such enrichment was not observed in chronic nonbacterial osteomyelitis. We also compared a previously reported core saliva microbiome (59 genera) with our data and found that out of the 74 genera detected in this study, 47 (including Mogibacterium) were not included in the previous meta-analysis. Additionally, we analyzed a core-genome tree of Mogibacterium from chronic bacterial osteomyelitis and healthy control samples along with a reference complete genome and found that Mogibacterium from both groups was indistinguishable at the core-genome and pan-genome levels. Although limited by the small sample size, our study provides novel evidence of a significant increase in Mogibacterium abundance in the chronic bacterial osteomyelitis group. Moreover, our study presents a comparative analysis of the taxonomic and sequence abundances of all genera detected using deep salivary shotgun metagenome data. The distinct enrichment of Mogibacterium suggests its potential as a marker to distinguish between patients with chronic nonbacterial osteomyelitis and chronic bacterial osteomyelitis, particularly at the early stages when differences are unclear.

Humans↗

An archaic reference-free method to jointly infer Neanderthal and Denisovan introgressed segments in modern human genomes.

Admixture between populations is a common feature of human history. Admixture events introduce new genetic variation that can fuel evolution. Characterizing the significance of admixture events on the evolution of populations across various species is of great interest to evolutionary geneticists. Local Ancestry Inference (LAI) methods infer genetic ancestry of an individual at a particular chromosomal location. Certain methods specialize in detecting archaic introgression, which consists of interbreeding between modern and archaic humans like Neanderthals and Denisovans. Most current LAI methods allow the detection of a single archaic ancestry, and post-processing may distinguish between multiple waves of introgression. These methods vary in how they choose archaic or modern reference genomes for the inference. Here, we present a new HMM-based method (DAIseg), which has the advantage of simultaneously distinguishing between multiple waves of ancient and recent admixture, using only modern human reference genomes. Simulations demonstrate that DAIseg achieves higher overall performance than state-of-the-art methods. We also apply DAIseg to Papuan populations to jointly detect Denisovan and Neanderthal introgressed segments, and identify a higher number of archaic segments than previous methods. Analysis of inferred introgressed segments, shows that we can identify evidence for two Denisovan introgression events in Papuans. Overall, on top of being able to deal with both Archaic and recent admixture, DAIseg provides a more principled approach for detecting and classifying Denisovan and Neanderthal segments which will improve downstream analysis of introgressed segments to infer the impact of archaic introgression in humans.

Denisovan↗

Detection of chromosomal gains and losses in comparative genomic hybridization analysis based on standard reference intervals.

Criteria for detection of chromosome aberrations by Comparative Genomic Hybridization (CGH) are not standardized and improvement of this part of the analysis is of paramount importance to the applicability of the technique. The aim of this work was to suggest CGH detection criteria that increase the specificity and sensitivity and at the same time include chromosome regions previously excluded from CGH analysis. We analyzed 33 hybridizations with normal DNA and modified our CGH software in order to use a selection of these normal analyses as a model for interpretation of analyses of unknown samples. This approach was successfully tested on 14 samples with known aberrations.

Chromosome Aberrations↗

A gridded genomic library of the honeybee (Apis mellifera): a reference library system for basic and comparative genetic studies of a hymenopteran genome.

We present a gridded genomic library of the honey-bee (Apis mellifera) for comparative and basic genetic study of the honeybee genome. The library will be established as a "Reference Library" system, and clones as well as data will be shared with the entire scientific community. This will accelerate the molecular level of honeybee genetics, combining the efforts of different laboratories. Because of male haploidy and the high rate of recombination, the honeybee is becoming a model organism for genomic studies of naturally occurring traits and behavioral genetics. The library consists of about 110,000 clones spotted at high density onto four filter membranes, representing 22 genome equivalents. Preliminary analysis using single-copy sequences revealed a positive clone number of the same order. The techniques for library generation and preliminary analysis as well as library access are described.

Animals↗

Genomic imprinting in testicular germ cell tumours.

Genomic imprinting refers to the parental origin-specific functional difference between the paternally and maternally-derived mammalian haploid genome. Normal embryogenesis depends on the presence of both a paternal and a maternal copy of particular chromosomal regions, containing the so-called imprinted genes. Genomic imprinting is established somewhere in the maturation from a primordial germ cell to a mature gamete, either spermatid or oocyte. We discuss the value of testicular cancers, especially those derived from the germ cell lineage, as a model to study erasement of the biparental pattern of genomic imprinting as present in the zygote and establishment of the paternal pattern during spermatogenesis. In addition, we will present data on the presence of X-inactivation in these cancers.

Animals↗

The inheritance of cognitive skills: does genomic imprinting play a role?

Genomic imprinting refers to the differential expression of a gene based on parental origin. Animal and clinical studies have suggested that genomic imprinting is influential in brain development, with the maternal genome playing a disproportionate role in the development of the cortex. The present study investigated this phenomenon in a nonclinical human population, using intrafamilial correlations. Broadly consistent with predictions, it was found that abilities mediated by frontal, parietal, and temporal lobes, but not occipital lobes, were more closely correlated between children and mothers versus fathers. The implications of these findings for the prevailing theory of the evolution of genomic imprinting, and for the general study of genetics and behavior, are discussed.

Adolescent↗

Theory of genomic imprinting conflict in social insects.

BACKGROUND: Genomic imprinting refers to the differential expression of genes inherited from the mother and father (matrigenes and patrigenes). The kinship theory of genomic imprinting treats parent-specific gene expression as products of within-genome conflict. Specifically, matrigenes and patrigenes will be in conflict over treatment of relatives to which they are differently related. Haplodiploid females have many such relatives, and social insects have many contexts in which they affect relatives, so haplodiploid social insects are prime candidates for tests of the kinship theory of imprinting. RESULTS: Matrigenic and patrigenic relatednesses are derived for individuals affected in a variety of contexts, including queen competition, sex ratio, worker laying of male eggs and policing, colony fission, and adoption of new queens. Numerous predictions emerge for what contexts should elicit imprinting, which individuals and tissues will show it, and the direction of imprinting effects. The predictions often vary for different genetic structures (varying queen and mate number) and often contrast with predictions for diploids. CONCLUSION: Because the contexts differ from the normal imprinting case, and because nothing is currently known about imprinting in social insects, these predictions can serve as a strong a priori test of the kinship theory of imprinting. If the predictions are correct, then social insects, which have long served as exemplars of cooperation between individuals, will also be shown to be extraordinary examples of competition within individual genomes.

Animals↗

Repeats mimic pathogen-associated patterns across a vast evolutionary landscape.

An emerging hallmark of many human diseases is transcription of typically silenced repetitive DNA containing pathogen-associated molecular patterns (PAMPs). These PAMPs engage the innate immune system via pattern recognition receptors (PRRs)-a phenomenon known as viral mimicry. We propose a statistical physics framework to quantify viral mimicry by measuring "selective forces" that enrich PAMPs compared to a genome-wide reference distribution. We validate our predictions by identifying repeats that bind different PRRs and show potential viral mimics in different repeat families across eukaryotic genomes, suggesting shared mechanisms drive emergence and retention. We propose two non-exclusive evolutionary hypotheses. The first "repeat-centric" hypothesis posits PAMPs are integral to the repeat life cycle and are therefore enriched as they mediate repeat expansion. The second "organism-centric" hypothesis proposes viral mimicry functions as a cell-intrinsic feedback mechanism for sensing and reacting to transcriptional dysregulation, which provides a selective pressure to maintain PAMPs in genomes.

Humans↗

DNABERT-S: pioneering species differentiation with species-aware DNA embeddings.

SUMMARY: We introduce DNABERT-S, a tailored genome model that develops species-aware embeddings to naturally cluster and segregate DNA sequences of different species in the embedding space. Differentiating species from genomic sequences (i.e. DNA and RNA) is vital yet challenging, since many real-world species remain uncharacterized, lacking known genomes for reference. Embedding-based methods are therefore used to differentiate species in an unsupervised manner. DNABERT-S builds upon a pre-trained genome foundation model named DNABERT-2. To encourage effective embeddings to error-prone long-read DNA sequences, we introduce Manifold Instance Mixup (MI-Mix), a contrastive objective that mixes the hidden representations of DNA sequences at randomly selected layers and trains the model to recognize and differentiate these mixed proportions at the output layer. We further enhance it with the proposed Curriculum Contrastive Learning (C2LR) strategy. Empirical results on 28 diverse datasets show DNABERT-S's effectiveness, especially in realistic label-scarce scenarios. For example, it identifies twice more species from a mixture of unlabeled genomic sequences, doubles the Adjusted Rand Index (ARI) in species clustering, and outperforms the top baseline's performance in 10-shot species classification with just a 2-shot training. AVAILABILITY AND IMPLEMENTATION: Model, codes, and data are publically available at https://github.com/MAGICS-LAB/DNABERT_S.

Sequence Analysis, DNA↗

DNABERT-S: Pioneering Species Differentiation with Species-Aware DNA Embeddings.

We introduce DNABERT-S, a tailored genome model that develops species-aware embeddings to naturally cluster and segregate DNA sequences of different species in the embedding space. Differentiating species from genomic sequences (i.e., DNA and RNA) is vital yet challenging, since many real-world species remain uncharacterized, lacking known genomes for reference. Embedding-based methods are therefore used to differentiate species in an unsupervised manner. DNABERT-S builds upon a pre-trained genome foundation model named DNABERT-2. To encourage effective embeddings to error-prone long-read DNA sequences, we introduce Manifold Instance Mixup (MI-Mix), a contrastive objective that mixes the hidden representations of DNA sequences at randomly selected layers and trains the model to recognize and differentiate these mixed proportions at the output layer. We further enhance it with the proposed Curriculum Contrastive Learning (C2LR) strategy. Empirical results on 23 diverse datasets show DNABERT-S's effectiveness, especially in realistic label-scarce scenarios. For example, it identifies twice more species from a mixture of unlabeled genomic sequences, doubles the Adjusted Rand Index (ARI) in species clustering, and outperforms the top baseline's performance in 10-shot species classification with just a 2-shot training. Model, codes, and data is publicly available at https://github.com/MAGlCS-LAB/DNABERT_S.

Journal Article↗

Genome duplications and other features in 12 Mb of DNA sequence from human chromosome 16p and 16q.

Several publicly funded large-scale sequencing efforts have been initiated with the goal of completing the first reference human genome sequence by the year 2005. Here we present the results of analysis of 11.8 Mb of genomic sequence from chromosome 16. The apparent gene density varies throughout the region, but the number of genes predicted (84) suggests that this is a gene-poor region. This result may also suggest that the total number of human genes is likely to be at the lower end of published estimates. One of the most interesting aspects of this region of the genome is the presence of highly homologous, recently duplicated tracts of sequence distributed throughout the p-arm. Such duplications have implications for mapping and gene analysis as well as the predisposition to recurrent chromosomal structural rearrangements associated with genetic disease.

Animals↗

Expanding vaginal microbiome pangenomes via a custom MIDAS database reveals Lactobacillus crispatus accessory genes associated with cervical dysplasia.

The vaginal microbiome plays a central role in reproductive health. Vaginal microbiome dysbiosis is associated with many adverse reproductive health outcomes, but most studies have focused on associations at the species level. The potential contribution of intraspecies microbial variation, especially gene content differences across bacterial strains, remains underexplored in reproductive health contexts. The Metagenomic Intra-Species Diversity Analysis (MIDAS) framework enables such analyses, but depends on comprehensive reference databases. We constructed a MIDAS-compatible pangenome database from over 18,000 genomes in the Vaginal Microbiome Genome Collection (VMGC). Compared to the Genome Taxonomy Database (GTDB)-derived reference, the VMGC-derived database expanded the pangenomes of prevalent vaginal species, better capturing vaginal-specific intraspecies diversity. Applying this database to vaginal samples from a cervical dysplasia cohort, we identified 13 Lactobacillus crispatus accessory genes significantly associated with cervical dysplasia, including a HicAB toxin-antitoxin system, three transcriptional regulators, and three phage-derived genes. These findings highlight the utility of body site-specific reference resources and shotgun metagenomic sequencing for uncovering intraspecies microbial variation relevant to reproductive health.IMPORTANCEThe vaginal microbiome plays a critical role in reproductive health, and different bacteria from the same species can carry different genes that influence how the strains interact with the host and other microbes. These strain-level differences are often overlooked when microbiomes are analyzed only at the species level. Existing genomic reference databases are heavily biased toward gut and environmental bacteria, leaving the genetic diversity of vaginal microbes understudied. We built a specialized reference database from over 18,000 vaginal bacterial genomes that better reflects this diversity. We then applied this resource to quantify gene-level variation in vaginal samples from a cervical dysplasia cohort. Focusing on Lactobacillus crispatus, a prevalent and often beneficial vaginal species, we identified 13 genes that were more common in women with cervical dysplasia than in controls. This work demonstrates that body site-specific genomic resources are essential for uncovering strain-level bacterial differences relevant to reproductive health.

Lactobacillus crispatus↗

The Retinome - defining a reference transcriptome of the adult mammalian retina/retinal pigment epithelium.

BACKGROUND: The mammalian retina is a valuable model system to study neuronal biology in health and disease. To obtain insight into intrinsic processes of the retina, great efforts are directed towards the identification and characterization of transcripts with functional relevance to this tissue. RESULTS: With the goal to assemble a first genome-wide reference transcriptome of the adult mammalian retina, referred to as the retinome, we have extracted 13,037 non-redundant annotated genes from nearly 500,000 published datasets on redundant retina/retinal pigment epithelium (RPE) transcripts. The data were generated from 27 independent studies employing a wide range of molecular and biocomputational approaches. Comparison to known retina-/RPE-specific pathways and established retinal gene networks suggest that the reference retinome may represent up to 90% of the retinal transcripts. We show that the distribution of retinal genes along the chromosomes is not random but exhibits a higher order organization closely following the previously observed clustering of genes with increased expression. CONCLUSION: The genome wide retinome map offers a rational basis for selecting suggestive candidate genes for hereditary as well as complex retinal diseases facilitating elaborate studies into normal and pathological pathways. To make this unique resource freely available we have built a database providing a query interface to the reference retinome 1.

Adult↗