Search PubMedSearch

SEARCH · Search PubMed

Results for “Genomic Structural Variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Molecular cloning and complete nucleotide sequence of the genome of Japanese encephalitis virus Beijing-1 strain.

The genomic RNA of the Japanese encephalitis virus (JEV) Beijing-1 strain was reversely transcribed and the synthesized cDNA was molecularly cloned. Six continuous cDNA clones that cover the entire virus genome were established and sequenced to determine the complete nucleotide sequence of the JEV RNA. The precise genomic size was estimated as 10,965 bases long. With flanking 95 bases at the 5' and 583 bases at the 3' non-coding regions, one long open reading frame (ORF) was revealed encoding a virus polyprotein with 3,429 amino acid residues. Because of sequence homologies observed between JEV and other flaviviruses, the genome organization of JEV appears to be identical with other flaviviruses. Genetic variation detected among flavivirus genomes is consistent with the established serological relatedness between JEV and other members of flaviviruses. The secondary structure of the JEV genome is deduced and discussed concerning its involvement in genome replication.

Base Sequence

Hepatitis E: review.

Hepatitis E is endemic, often provoking epidemics in many developing countries. It resembles hepatitis A clinically and epidemiologically but show a higher mortality rate and less infectiousness. Several lines of evidence strongly support the assumption that humans become immunized once they contract hepatitis E. Because of the low infectiousness, most of the adult population of endemic areas are susceptible to hepatitis E until an epidemic occurs, although they are almost always infected with hepatitis A during infancy. Epidemics are caused by accidental contamination by the hepatitis E virus (HEV) in feces of water provided to these people. The liver change reveals necroinflammation related to the immune-mediated mechanism. The HEV is molecularly cloned and sequenced and has a single-stranded, positive-sense RNA genome, 7,194 nucleotides followed by a poly (A) tail. There are three open reading frames. The non-structural gene, approximately 5 kb is located at the 5' end, while the structural gene, approximately 2 kb is located at the 3' end of the genome. There is a low level of nucleotide variations among HEV strains isolated from Myanmar and China and a single serotype appears to exist. The HEV may be a new RNA virus or belong to Caliciviridae family. Further investigation include in vitro propagation, elucidation of the gene replication, global seroepidemiology and vaccination of the HEV.

Animals

Molecular cloning and nucleotide sequence of a variant wheat histone H4 gene.

To determine whether there is structural variation among histone H4 genes in wheat, one (TH091) of the H4 genes that had been cloned from a wheat genomic DNA library was sequenced and compared with another H4 gene (TH011) which we had described previously [Tabata et al., Nucl. Acids Res. 11 (1983) 5865-5865]. Nucleotide sequence analysis revealed that there are 17 nucleotide replacements in the protein-coding region of two H4 genes, causing only one amino acid substitution: a glycine at position 4 (from the N terminus) in TH011 was replaced by an aspartic acid in TH091. S1 mapping, using total nuclear RNA from germinated seeds, indicated that the H4 gene was transcribed in vivo.

Amino Acid Sequence

Direct sequencing of large flavivirus PCR products for analysis of genome variation and molecular epidemiological investigations.

The polymerase chain reaction (PCR) was used to amplify viral cDNAs from selected regions of dengue genomic RNA by using appropriate 'consensus' primers. DNA amplicons containing the structural genes from all 4 dengue serotypes were prepared and directly sequenced using dengue-virus-specific primers. This method can characterize reliably flavivirus field isolates at the molecular level without extensive virus propagation and molecular cloning, and will be a valuable tool for molecular epidemiological studies.

Base Sequence

The current and future perspective of ChickenGTEx project and its applications in precision breeding.

The Chicken Genotype-Tissue Expression (ChickenGTEx) project was established to systematically characterize the regulatory landscape of the chicken genome and to accelerate the translation of functional genomics into precision breeding. By integrating whole-genome sequencing with multi-tissue transcriptomic profiling, ChickenGTEx provides a comprehensive atlas of gene expression regulation across diverse tissues and physiological systems. Current findings demonstrate that complex production traits are governed by coordinated regulatory networks rather than isolated loci, with substantial contributions from tissue-specific gene expression, structural variation, and genotype-by-sex interactions. Sex-dependent regulatory effects further refine the genetic architecture of metabolic, immune, and reproductive traits, highlighting the importance of incorporating sex as a biological variable in genomic analyses. Application of integrative omics frameworks within elite layer populations has revealed multilayer regulatory mechanisms underlying extended laying performance, feed efficiency, metabolic health, and eggshell quality. By partitioning phenotypic variance into genetic, regulatory, and host-microbiome components, these approaches move beyond association-based mapping toward causal inference and biological interpretation. Importantly, validated regulatory loci identified through ChickenGTEx and related analyses provide actionable markers for genomic selection and rational targets for precision genome modification. Looking forward, continued expansion of regulatory atlases, incorporation of single-cell and longitudinal data in diverse environmental conditions, and integration of functional annotation into breeding pipelines will further enhance prediction accuracy and sustainable genetic improvement. The ChickenGTEx project thus represents a foundational platform bridging functional genomics and practical poultry breeding.

Animals

Extensive allelic variation in Cryptococcus neoformans.

The orotidine monophosphate pyrophosphorylase (OMPPase) gene locus of the DNA of 13 Cryptococcus neoformans var. neoformans strains, including 10 recent clinical isolates, was studied by using restriction fragment length polymorphisms and nucleotide sequence analysis. The OMPPase locus (URA5) is highly polymorphic, and at least six alleles were identified. The nucleotide sequences of some alleles differed by up to 5%. The majority of the nucleotide polymorphisms in the protein-coding region occurred at the third codon position and were silent. The low frequency of replacement nucleotide substitutions relative to silent nucleotide substitutions implied that there is strong selection against amino acid changes in OMPPase. The allelic variation suggested that there is extensive genomic diversity among C. neoformans clinical isolates from one geographic area. The various alleles are potentially useful markers in the study of the population structure, epidemiology, and pathogenesis of C. neoformans strains.

Alleles

[Genetic diversity analysis of Forsythia suspensa germplasm resources in Shanxi based on phenotypic traits and SNP molecular markers].

This study aimed to clarify the degree of fruit phenotypic variation and the characteristics of genetic diversity, population structure, and genetic differentiation of Forsythia suspensa resources in Shanxi, providing an important basis for germplasm conservation and breeding of superior varieties. A total of 46 F. suspensa fruits were collected, and 12 agronomic traits were measured and analyzed. The population genetic structure and genetic diversity of F. suspensa germplasm were evaluated using simplified genome sequencing technology. For the five quality traits of the 46 fruits, the Shannon-Wiener index ranged from 0.631 to 1.074, and the Simpson index ranged from 0.379 to 0.560. The seven quantitative traits exhibited abundant genetic variation, with coefficients of variation ranging from 9.764%(fruit shape index) to 45.494%(forsythin content). Principal component analysis reduced the 12 phenotypic traits to four factors, with a cumulative variance contribution of 74.547%. Sequencing data showed mean Q20 and Q30 values of 98.13% and 94.33%, respectively, with an average GC content of 35.95%. After filtering, a total of 12 347 327 high-quality single nucleotide polymorphism(SNP) loci were obtained. Based on these high-quality SNPs, principal component analysis, population structure analysis, and phylogenetic tree construction were carried out. The 46 germplasm resources were divided into four groups; however, grouping showed little relationship with geographic origin, and intermixing occurred among regions. Mantel test revealed a significant but weak positive correlation between phenotypic and genetic distances(r=0.159, P=0.001). At the molecular level, the four groups exhibited moderate genetic diversity overall, and the genetic differentiation index among populations ranged from 0.027 to 0.084, indicating low to moderate differentiation. The rich genetic diversity of the main phenotypic traits provides a solid material basis for screening superior germplasm and genetic breeding of F. suspensa.

Forsythia

High spontaneous mutation rate of Rous sarcoma virus demonstrated by direct sequencing of the RNA genome.

Direct and extensive sequencing of RSV RNA genome is reported. More than 10,000 nt of the T1 RNase resistant RSV RNA fragments (1) have been sequenced and shown to cover 3900 nt of RSV genome. The frequent sequence variations found indicate that RSV supports a very high incidence of spontaneous mutations in the course of replication, one very probable cause of the genetic diversity among the avian retroviruses. Sequences of the structured RSV RNAs allowed us also to precisely characterize the structured domains of the retroviral genome and show that the src gene is not structured.

Amino Acid Sequence

Minisatellite variant repeat (MVR) mapping: analysis of 'null' repeat units at D1S8.

Minisatellite variant repeat mapping by PCR (MVR-PCR) is a new approach to studying variation in human DNA which analyses interspersion patterns of variant repeats within minisatellite arrays. MVR-PCR has been applied to the hypervariable human minisatellite D1S8 which contains two major classes of variant 29bp repeat units designated a-type and t-type. The MVR-PCR assay uses a- or t-type specific primers, together with an amplimer at a fixed site in the DNA flanking the minisatellite, to reveal the interspersion patterns of variant repeats along an allele. Extreme levels of variation are seen both in the internal structures of individual alleles and in the digital code generated from the two superimposed alleles in total genomic DNA. However, occasional repeat units fail to amplify in MVR-PCR, signifying the existence of further repeat sequence variants termed 'null' or O-type repeats. Although not significant in individual identification, correct genotyping of null repeats is important when using MVR digital codes in parentage analysis. We have therefore characterised these null repeats and show that most null repeats share a common variant repeat sequence. We discuss the possible origins of null repeats and their application to paternity testing and the analysis of minisatellite evolution.

Alleles

An integrated human immunoglobulin germline resource linking allele diversity to expressed repertoire structure.

Human immunoglobulin (IG) loci are highly polymorphic, yet existing germline resources remain noisy and incomplete, limiting our ability to link inherited variation to antibody repertoires. Here, we integrate high-fidelity long-read genomic sequencing with matched adaptive immune receptor repertoire sequencing (AIRR-seq) to construct HUSA, a population-scale, evidence-resolved germline resource. Using a conservative allele inference framework, HUSA expands current references more than three-fold, identifying over 1300 alleles while preserving allele-level evidence provenance across genomic and repertoire data. By linking genotype and expressed repertoires within individuals, we show that coding-region similarity predicts the structure of adjacent recombination signal sequences and leader regions, revealing that IG alleles are organized as linked cis-regulatory units associated with differences in recombination context and allele usage. These results define key germline constraints shaping repertoire formation and establish a robust, genotype-aware foundation for the analysis of immune receptor repertoires.

Journal Article

Genomic-based revelation of genetic structure and adaptive characterization of Schizopygopsis malacanthus in the Jinsha River and Yalong River.

BACKGROUND: As a highly specialized class of schizothoracine fishes, Schizopygopsis malacanthus has attracted much attention due to its widespread distribution. To investigate the impact of the Qinghai‒Tibet movement on S. malacanthus, we analyzed the genetic evolutionary history of this species. RESULTS: These results showed that there was a high level of genetic differentiation between Jinsha River (JSR) populations and Yalong River (YLR) populations. The genetic diversity of intra-YLR populations was higher than that of the intra-JSR populations. There was gene exchange of the Suwalong population to the Huoqu and Ganzi populations. Furthermore, both of the JSR and YLR populations exhibited a gradual increase in the genetic differentiation index from low to high altitudes, and the effective population of high-elevation populations has gradually expanded. In high-altitude populations, the selected genes were enriched in DNA repair, light transduction, and energy metabolism, reflecting the genetic basis for their migration to higher altitudes. CONCLUSIONS: S. malacanthus populations had the higher genetic differentiation and genetic diversity in the JSR and its main tributary YLR. Therefore, we should preserve high-elevation natural river sections as much as possible and reserve habitats for their migration and diffusion.

Animals

Sequence analysis of 22 kDa-like alpha-coixin genes and their comparison with homologous zein and kafirin genes reveals highly conserved protein structure and regulatory elements.

Several genomic and cDNA clones encoding the 22 kDa-like alpha-coixin, the alpha-prolamin of Coix seeds, were isolated and sequenced. Three contiguous 22 kDa-like alpha-coixin genes designated alpha-3A, alpha-3B and alpha-3C were found in the 15 kb alpha-3 genomic clone. The alpha-3A and alpha-3C genes presented in-frame stop codons at position +652. The two genes with truncated ORFs are flanking the alpha-3B gene, suggesting that the three alpha-coixin genes may have arisen by tandem duplication and that the stop codon was introduced before the duplication. Comparison of the deduced amino acid sequences of alpha-coixin clones with the published sequences of 22 kDa alpha-zein and 22 kDa-like alpha-kafirin revealed a highly conserved protein structure. The protein consists of an N-terminus, containing the signal peptide, followed by ten highly conserved tandem repeats of 15-20 amino acids flanked by polyglutamines, and a short C-terminus. The difference between the 22 kDa-like alpha-prolamins and the 19 kDa alpha-zein lies in the fact that the 19 kDa protein is exactly one repeat motif shorter than the 22 kDa proteins. Several putative regulatory sequences common to the zein and kafirin genes were identified within both the 5' and 3' flanking regions of alpha-3B. Nucleotide sequences that match the consensus TATA, CATC and the ca. -300 prolamin box are present at conserved positions in alpha-3B relative to zein and kafirin genes. Two putative Opaque-2 boxes are present in alpha-3B that occupies approximately the same positions as those identified for the 22 kDa alpha-zein and alpha-kafirin genes. Southern hybridization, using a fragment of a maize Opaque-2 cDNA clone as a probe, confirmed the presence of Opaque-2 homologous sequences in the Coix and sorghum genomes. The overall results suggest that the structural and regulatory genes involved in the expression of the 22 kDa-like alpha-prolamin genes of Coix, sorghum and maize, originated from a common ancestor, and that variations were introduced in the structural and regulatory sequences after species separation.

Amino Acid Sequence

The structure of hepatitis B envelope and molecular variants of hepatitis B virus.

Accumulated evidence in recent years has shown that the variation of hepatitis B virus (HBV) genomes may have profound implications for our understanding of hepatitis B pathogenesis and prevention. Attention has focused on areas of the outer envelope coded by the S gene which are involved in the induction of a protective neutralising antibody response, and mutations which directly affect the production of C gene products, one of which is considered as a target for immune T cells involved in virus clearance. This review highlights recent experimental data which emphasizes the role of such mutations in the establishment and maintenance of chronic HBV infections and focuses attention on the significance of HBV variants with respect to the expanding use of HBV vaccines for mass immunization.

Amino Acid Sequence

Structure of the intergenic spacer region from the ribosomal RNA gene family of white spruce (Picea glauca).

Five genomic clones containing ribosomal DNA repeats from the gymnosperm white spruce (Picea glauca) have been isolated and characterized by restriction enzyme analysis. No nucleotide variation or length variation was detected within the region encoding the ribosomal RNAs. Four clones which contained the intergenic spacer (IGS) region from different rDNA repeats were further characterized to reveal the sub-repeat structure within the IGS. The sub-repeats were unusually long, ranging from 540 to 990 bp but in all other respects the structure of the IGS was very similar to the organization of the IGS from wheat, Drosophila and Xenopus.

DNA, Ribosomal

Scaling up orphan crop research: genebank genetics highlight geographic structure in cultivated cowpea from 10 617 global accessions.

Vigna unguiculata (L.) Walp. is a dryland legume crop, providing essential food and nutritional security for millions of people across the semi-arid tropics, in Africa, Asia and Latin America. However, as a typical 'orphan crop', cowpea has long remained underrepresented in global genomic research to support crop improvement. Here, we conducted the largest genetic diversity analysis of cowpea to date, comprising 10 617 accessions sourced from seven international collections. Using genotyping-by-sequencing, we characterised the global patterns of genetic diversity, assessed redundancy within and across collections, and examined the geographic structure of the cowpea global allele pool. Our results revealed nine distinct genetic groups with clear geographic associations and fine-scale population differentiation, reflecting dispersal history, regional adaptation and the influence of modern breeding. Duplication across collections was detected, highlighting the need for improved curation and integration of germplasm resources. Landraces from sub-Saharan Africa do not fully capture the genetic diversity present in several other geographic regions, indicating the existence of abundant and untapped genetic resources worldwide. These findings not only provide insights into the genetic structure and evolutionary history of cowpea but also offer a valuable foundation for harnessing global germplasm diversity to enhance breeding potential and accelerate crop improvement.

Vigna

Proteomic comparison of epidemic Australian Bordetella pertussis biofilm cells.

Bordetella pertussis causes whooping cough, a severe respiratory infectious disease. Studies have compared the currently dominant single nucleotide polymorphism (SNP) cluster I (pertussis toxin promoter allele, ptxP3) and previously dominant SNP cluster II (ptxP1) strains as planktonic cells. Since biofilm formation is linked with B. pertussis pathogenesis in vivo, this study compared the biofilm formation capabilities of representative strains of cluster I and cluster II. Confocal laser scanning microscopy found that the cluster I strain had a denser biofilm structure compared to the cluster II strain. Differences in protein abundance of the biofilm cells were then compared using tandem mass tagging and high-resolution multiple reaction monitoring. In total, 1,453 proteins were identified, of which 40 proteins had significant differential abundance between the two strains in biofilm conditions. Of particular interest was a large increase in the abundance of energy metabolism proteins (cytochrome proteins PetABC and BP3650) in the cluster I strain. When the abundance of these proteins was compared between six additional strains from each cluster, it was found that the protein abundance varied between all strains. These findings suggest that there are large levels of individual proteomic diversity between B. pertussis strains in biofilm conditions despite the highly conserved genome of the species. Overall, this study revealed visual differences in biofilm structure between B. pertussis strains and highlighted strain-specific variation in protein abundance that dominates potential cluster-specific changes that may be linked with the dominance of cluster I strains.IMPORTANCEBordetella pertussis causes whooping cough. The currently circulating cluster I strains have taken over previously dominant cluster II strains. It is important to understand the reasons behind this evolution to develop new strategies against the pathogen. Recent studies have shown that B. pertussis can form biofilms during infection. This study compared the biofilm formation capabilities of a cluster I and a cluster II strain and identified visual differences in the biofilms. The protein abundance between these strains grown in biofilms was compared, and proteins identified with varied abundance were measured with additional strains from each cluster. It was found that despite the highly conserved genetics of the species, there was varied protein abundance between the additional strains. This study highlights that strain-specific variation in protein abundance during biofilm conditions may dominate the cluster-specific changes that may be linked to the dominance of cluster I strains.

Bordetella pertussis