Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequencing quality”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Mobile genetic elements-driven partitions of mega-plasmids resistome in Salmonella Infantis.

Salmonella enterica serovar Infantis (S. Infantis) becomes the primary pathogen among the top Salmonella serotypes, contributing to numerous cases of foodborne illness annually in the United States. S. Infantis infection has spread rapidly worldwide, especially the clones with pESI-like plasmids. However, the underlying mechanisms regarding the transmission of S. Infantis, particularly mobile genetic elements (MGEs), mediated horizontal gene transfer, are limited. The objective of this study was to evaluate the relationship, if any, among MGEs, antibiotic-resistant genes (ARGs), and virulence factors (VFs) within S. Infantis via genomic analysis. A total of 91 S. Infantis complete genomes with high sequencing quality were selected for downstream bioinformatic analysis. The results showed that the majority of VFs were located in the bacterial chromosomes, while most ARGs were carried by S. Infantis mega-plasmids in an MGE-favored manner. Integrons and transposons were closely associated with certain ARGs, but prophages within mega-plasmids displayed a diverse ARG profile. Collectively, MGE-mediated horizontal gene transfer might lead to ARG acquisition by mega-plasmids, subsequently contributing to the resistome of S. Infantis. Our findings provide insights into the development of MGE-associated resistome in S. Infantis that could inform more effective prevention and intervention strategies to control this pathogen, further ensuring public health and safety.IMPORTANCEThe rapid emergence and transmission of antibiotic-resistant foodborne pathogens pose a significant risk to public health, necessitating the discovery of underlying mechanisms to control multidrug-resistant pathogens. Salmonella enterica serovar Infantis (S. Infantis) has become a pathogen of clinical and epidemiological relevance in recent years, ranking as the top prevalent serovar associated with foodborne illnesses and exhibiting resistance to several antibiotics. The current investigation of multidrug resistance (MDR) S. Infantis strains primarily emphasized the presence of mega-plasmids. However, the question of how mega-plasmids contribute to the transmission of antibiotic-resistant genes (ARG) is unaddressed. Utilizing the genomic characterization of S. Infantis complete genomes with high quality, our study revealed that the resistome of S. Infantis mega-plasmids-the primary ARG reservoirs of S. Infantis-followed a specific pattern of mobile genetic elements (MGEs). Monitoring the spread of MGE-carried ARGs within mega-plasmids should be considered in future surveillance.

Interspersed Repetitive Sequences↗

Optimization of a duplex amplification and sequencing strategy for the HVI/HVII regions of human mitochondrial DNA for forensic casework.

A duplex primer set for the amplification of mitochondrial DNA HVI and HVII control regions was evaluated for the optimization of a DNA sequencing protocol suitable for forensic casework. HVI and HVII products, with the absence of non-specific products, could be detected by agarose gel electrophoresis when as little as 0.5 and 0.1pg of DNA were amplified for 34 and 38 cycles, respectively. Because HVI and HVII amplicons are co-synthesized in the duplex PCR, fewer steps are required (lessening the risk of cross contamination events) and more frugal use of precious extracted DNA samples is possible, both desirable features for forensic casework. The ABI Prism BigDyetrade mark version 1.1 chemistry provided high quality sequencing data, with little or no background noise and uniform peak heights, outcomes that favored reliable detection of heteroplasmy, particularly at early sequence reads (<40 bases). Optimal compromise between sensitivity and sequence accuracy in the absence of noise was achieved starting at 150 mitochondrial genome copies. The protocol is effective (no sequence errors) with highly degraded DNA (average detectable template size of 200bp). Dual artificial template mixtures with the minor component at 15% suggests that heteroplasmy should be detected at this level with confidence.

Complementarity Determining Regions↗

Nucleotide substitutions in Staphylococcus aureus strains, Mu50, Mu3, and N315.

A specific phenotype of Staphylococcus aureus strains Mu50 and Mu3 is characterized by thickened cell wall and moderate resistance to vancomycin. The N315 strain is a prototype of methicillin-resistant S. aureus (MRSA), but it is methicillin susceptible, despite carrying the mecA resistance gene. Here, we revised differences in the sequences of Mu50 and N315, referencing that of Mu3 which were assumed to be of one lineage. The 362 ORFs diverse between Mu50 and N315 were picked up, and the corresponding ones in three strains were re-sequenced. This defined 213 ORFs diverse between Mu50 and N315, and 9 between Mu50 and Mu3. The fixed diversities of 174 ORFs (except for 39 silent ORFs from 213), including nucleotide substitution (NSs), frame shift, and truncation were grouped into three major functional categories, which were transport (14.9% in the 174 diverse ORFs), metabolism of carbohydrates (5.7%), and RNA synthesis (9.6%). The other gene categories had small diversities. These gene categories seemed to be functionally decisive for the Mu50-specific characters, the thickened cell wall and moderate vancomycin resistance. All of the diverse genes and the high quality sequence of Mu50 can be viewed at the web site (http://133.5.48.239/VRSA/).

Bacterial Proteins↗

The evolution of retrotransposon regulatory regions and its consequences on the Drosophila melanogaster and Homo sapiens host genomes.

It has now been established that transposable elements (TEs) make up a variable, but significant proportion of the genomes of all organisms, from Bacteria to Vertebrates. However, in addition to their quantitative importance, there is increasing evidence that TEs also play a functional role within the genome. In particular, TE regulatory regions can be viewed as a large pool of potential promoter sequences for host genes. Studying the evolution of regulatory region of TEs in different genomic contexts is therefore a fundamental aspect of understanding how a genome works. In this paper, we first briefly describe what is currently known about the regulation of TE copy number and activity in genomes, and then focus on TE regulatory regions and their evolution. We restrict ourselves to retrotransposons, which are the most abundant class of eukaryotic TEs, and analyze their evolution and the subsequent consequences for host genomes. Particular attention is paid to much-studied representatives of the Vertebrates and Invertebrates, Homo sapiens and Drosophila melanogaster, respectively, for which high quality sequenced genomes are available.

Animals↗

Refined annotation of the Arabidopsis genome by complete expressed sequence tag mapping.

Expressed sequence tags (ESTs) currently encompass more entries in the public databases than any other form of sequence data. Thus, EST data sets provide a vast resource for gene identification and expression profiling. We have mapped the complete set of 176,915 publicly available Arabidopsis EST sequences onto the Arabidopsis genome using GeneSeqer, a spliced alignment program incorporating sequence similarity and splice site scoring. About 96% of the available ESTs could be properly aligned with a genomic locus, with the remaining ESTs deriving from organelle genomes and non-Arabidopsis sources or displaying insufficient sequence quality for alignment. The mapping provides verified sets of EST clusters for evaluation of EST clustering programs. Analysis of the spliced alignments suggests corrections to current gene structure annotation and provides examples of alternative and non-canonical pre-mRNA splicing. All results of this study were parsed into a database and are accessible via a flexible Web interface at http://www.plantgdb.org/AtGDB/.

Alternative Splicing↗

Single nucleotide polymorphism (SNP) discovery in duplicated genomes: intron-primed exon-crossing (IPEC) as a strategy for avoiding amplification of duplicated loci in Atlantic salmon (Salmo salar) and other salmonid fishes.

BACKGROUND: Single nucleotide polymorphisms (SNPs) represent the most abundant type of DNA variation in the vertebrate genome, and their applications as genetic markers in numerous studies of molecular ecology and conservation of natural populations are emerging. Recent large-scale sequencing projects in several fish species have provided a vast amount of data in public databases, which can be utilized in novel SNP discovery in salmonids. However, the suggested duplicated nature of the salmonid genome may hamper SNP characterization if the primers designed in conserved gene regions amplify multiple loci. RESULTS: Here we introduce a new intron-primed exon-crossing (IPEC) method in an attempt to overcome this duplication problem, and also evaluate different priming methods for SNP discovery in Atlantic salmon (Salmo salar) and other salmonids. A total of 69 loci with differing priming strategies were screened in S. salar, and 27 of these produced approximately 13 kb of high-quality sequence data consisting of 19 SNPs or indels (one per 680 bp). The SNP frequency and the overall nucleotide diversity (3.99 x 10-4) in S. salar was lower than reported in a majority of other organisms, which may suggest a relative young population history for Atlantic salmon. A subset of primers used in cross-species analyses revealed considerable variation in the SNP frequencies and nucleotide diversities in other salmonids. CONCLUSION: Sequencing success was significantly higher with the new IPEC primers; thus the total number of loci to screen in order to identify one potential polymorphic site was six times less with this new strategy. Given that duplication may hamper SNP discovery in some species, the IPEC method reported here is an alternative way of identifying novel polymorphisms in such cases.

Animals↗

Computational comparison of two mouse draft genomes and the human golden path.

BACKGROUND: The availability of both mouse and human draft genomes has marked the beginning of a new era of comparative mammalian genomics. The two available mouse genome assemblies, from the public mouse genome sequencing consortium and Celera Genomics, were obtained using different clone libraries and different assembly methods. RESULTS: We present here a critical comparison of the two latest mouse genome assemblies. The utility of the combined genomes is further demonstrated by comparing them with the human 'golden path' and through a subsequent analysis of a resulting conserved sequence element (CSE) database, which allows us to identify over 6,000 potential novel genes and to derive independent estimates of the number of human protein-coding genes. CONCLUSION: The Celera and public mouse assemblies differ in about 10% of the mouse genome. Each assembly has advantages over the other: Celera has higher accuracy in base-pairs and overall higher coverage of the genome; the public assembly, however, has higher sequence quality in some newly finished bacterial artificial chromosome clone (BAC) regions and the data are freely accessible. Perhaps most important, by combining both assemblies, we can get a better annotation of the human genome; in particular, we can obtain the most complete set of CSEs, one third of which are related to known genes and some others are related to other functional genomic regions. More than half the CSEs are of unknown function. From the CSEs, we estimate the total number of human protein-coding genes to be about 40,000. This searchable publicly available online CSEdb will expedite new discoveries through comparative genomics.

Animals↗

FunnyBase: a systems level functional annotation of Fundulus ESTs for the analysis of gene expression.

BACKGROUND: While studies of non-model organisms are critical for many research areas, such as evolution, development, and environmental biology, they present particular challenges for both experimental and computational genomic level research. Resources such as mass-produced microarrays and the computational tools linking these data to functional annotation at the system and pathway level are rarely available for non-model species. This type of "systems-level" analysis is critical to the understanding of patterns of gene expression that underlie biological processes. RESULTS: We describe a bioinformatics pipeline known as FunnyBase that has been used to store, annotate, and analyze 40,363 expressed sequence tags (ESTs) from the heart and liver of the fish, Fundulus heteroclitus. Primary annotations based on sequence similarity are linked to networks of systematic annotation in Gene Ontology (GO) and the Kyoto Encyclopedia of Genes and Genomes (KEGG) and can be queried and computationally utilized in downstream analyses. Steps are taken to ensure that the annotation is self-consistent and that the structure of GO is used to identify higher level functions that may not be annotated directly. An integrated framework for cDNA library production, sequencing, quality control, expression data generation, and systems-level analysis is presented and utilized. In a case study, a set of genes, that had statistically significant regression between gene expression levels and environmental temperature along the Atlantic Coast, shows a statistically significant (P < 0.001) enrichment in genes associated with amine metabolism. CONCLUSION: The methods described have application for functional genomics studies, particularly among non-model organisms. The web interface for FunnyBase can be accessed at http://genomics.rsmas.miami.edu/funnybase/super_craw4/. Data and source code are available by request at jpaschall@bioinfobase.umkc.edu.

Animals↗

Phylogenetic inconsistency of pairwise SNP clustering for inferring tuberculosis transmission in a high-burden, endemic setting: a case study from Thailand.

Whole-genome sequence analysis is now widely used to delineate tuberculosis transmission clusters. A standard practice is to cluster bacterial isolates based on a fixed maximum genome-wide pairwise single nucleotide polymorphism (pwSNP) distance threshold. In this study, we evaluated the phylogenetic consistency of pwSNP-distance clustering with thresholds ranging between 1 and 25 single nucleotide polymorphisms (SNPs) using two contrasting data sets: (i) a data set from the UK (N = 390) published by T. M. Walker, C. L. C. Ip, R. H. Harrell, J. T. Evans, et al. (Lancet Infect Dis 13:137-146, 2013, https://doi.org/10.1016/S1473-3099(12)70277-3), which was foundational to the establishment of this method, and (ii) a data set from Thailand (N = 3,341), characterized by persistent transmission and sparse, non-systematic sampling. For the UK data set, the standard pwSNP-distance clustering using thresholds of &#x2265;12 SNPs yielded entirely monophyletic clusters and showed high concordance with a comparative monophyly constrained, tree-based method. In contrast, for the Thai data set, pwSNP-distance clustering often generated non-monophyletic clusters, even by the 25-SNP threshold. The pwSNP-distance and comparative tree-based clustering methods only showed large consistency at thresholds of &#x2265;22 SNPs. This suggests that SNP clusters defined by low distance thresholds (i.e., <12 SNPs for the UK data set, and <22 SNPs for the Thai data set) may lack robustness, and the problem is particularly severe for data sets characterized by persistent transmission, likely due to poorer cluster separation. Moreover, our findings indicate that large cluster sizes, high maximum intra-cluster genetic distances, and broad sample collection time spans may serve as useful indicators of potentially non-monophyletic clusters. We also demonstrate that mixed infections can produce spurious, phylogenetically long-range SNP linkages, underscoring the necessity of strict sequence quality control.IMPORTANCEFixed-threshold pairwise single nucleotide polymorphism (pwSNP)-distance clustering is commonly used to delineate tuberculosis transmission clusters. From an epidemiological perspective, a genuine transmission cluster must be monophyletic, originating from a single source. However, pwSNP-distance clustering is inherently simplistic and can therefore violate this principle, making the assessment of its phylogenetic consistency critical. Our results demonstrate that while this method effectively delineated complete transmission clusters for the data set from the UK, a low-burden and non-persistent transmission setting, it frequently generated non-monophyletic clusters when applied to the Thai data set, characterized by persistent transmission alongside sparse and non-systematic sampling. Furthermore, we found that clusters derived using low distance thresholds could notably vary between the pwSNP-distance and comparative tree-based clustering methods, suggesting limited reliability and robustness. To accurately delineate tuberculosis transmission clusters, especially for complex data from high-burden, endemic settings, we recommend transitioning from pwSNP-distance clustering toward more robust, phylogenetic clustering that respects evolutionary descent.

Mycobacterium tuberculosis↗

An efficient, automatable template preparation for high throughput sequencing.

We have developed a 96-well format for DNA template isolation that can be readily automatable. The template isolation protocol involves simple alkaline lysis chemistry and reversible capture on a silica solid phase. After the cells are lysed, no centrifugation is necessary, as lysate purification, DNA binding, washing, and release occur in 96-well filter plates. Large numbers of templates prepared using the silica purification method have been sequenced and analyzed. The quality of sequence resulting from our method has been compared with that generated from several commercial plasmid preparation protocols. We found sequence quality of the silica bead preparations to be equivalent to or, in some cases, better than those prepared by other methods. This method offers many advantages over other protocols we have used. First, the silica purifications have allowed us to more than double overall laboratory throughput while decreasing our template isolation materials cost at least five-fold. Second, because we have eliminated all centrifugation steps in the protocol, automation has been much simpler. The protocol has also been adapted to purify PCR products for use as templates in subsequent sequencing reactions.

Automation↗

Accuracy of T1 measurement in dynamic contrast-enhanced breast MRI using two- and three-dimensional variable flip angle fast low-angle shot.

In vivo T1 measurements, used to monitor the uptake of contrast agent by tissues, are typically performed as a first step in implementing compartmental analysis of contrast-enhanced breast magnetic resonance imaging (MRI) data. We have extended previously described methodology for in vivo T1 measurement (using a variable flip-angle gradient-recalled echo technique) to two-dimensional (2D), fast low-angle shot (FLASH). This approach requires computational modeling of slice-selective radiofrequency (RF) excitation to correct for nonrectangular slice profiles. The accuracy with which breast tissue T1 values can be measured by this approach is examined: T1 measurements from phantom and in vivo image data acquired with 2D and 3D FLASH imaging sequences are presented. Significant sources of error due to imaging pulse sequence quality and RF transmit field nonuniformity in the breast coil device that will have detrimental consequences for compartmental analysis are identified. Rigorous quality assurance programs with calibrated phantoms are thus recommended, to verify the accuracy with which T1 measurements are obtained.

Breast↗

MRI of the knee: value of short echo time fast spin-echo using high performance gradients versus conventional spin-echo imaging for the detection of meniscal tears.

OBJECTIVE: Fast spin-echo (FSE) sequences reduce imaging time compared with conventional spin-echo (CSE) sequences, but may result in blurring. High-performance gradients permit shorter interecho spacing and use of the second echo as the effective TE (20 ms); both improvements reduce blurring. This randomized observer study compared a short TE, second-echo FSE sequence obtained using high-performance gradients and a CSE sequence with similar TR/TE for the detection of meniscal tears in the knee. DESIGN AND PATIENTS: One hundred consecutive MR examinations of the knee using FSE and CSE sequences at 1.5 T were evaluated. The FSE sequence used an effective TE of 20 ms (centered on the second echo at 2 times minimal interecho spacing) and an echo train length of 4. FSE and CSE parameters were otherwise similar. Four independent, masked readers reviewed randomized sagittal FSE and CSE sequences. RESULTS: Cases were assessed for the presence or absence of meniscal tears and, if present, whether tears were medial or lateral and anterior or posterior. Sequence concordance was 93.5% (1496 of 1600 meniscal segments); the intermethod kappa value was 0.78. Sequence quality was graded from 1 to 5. Average quality of CSE images was slightly but statistically significantly preferred by three of the four readers. CONCLUSION: There was no statistically significant difference between CSE imaging and FSE imaging centered on the second echo (20 ms) using high-performance gradients for the detection of meniscal tears in the knee. There was a small preference for the quality of CSE images.

Adolescent↗

Developing a simplified river landscape assessment model: examples from the Chungkang and Touchien rivers, Taiwan.

Currently, river landscape evaluations cannot be conducted by the general public, due to their lack of professional training. However, consulting professionals is time consuming and costly. The research conducted addresses both problems by (1) developing suitable criteria for assessing river environments, and (2) formulating a strategy for using the proposed criteria, thereby creating an effective method for river management by non-professionals. This research was carried out, in accordance with Visual Resource Management theory, at 12 survey sites along the ChungKang River. The landscape quality sequences acquired were then evaluated using a revised and simplified assessment model. The same process was repeated on the Touchien River to verify its feasibility. This research developed both specific criteria as well a method for evaluating river landscapes, which can be employed by non-professional river project managers. Ultimately, the aim of this research is to develop and promote sustainable river resource management.

Environmental Health↗

Multiple group-specific sequencing primers for reliable and rapid DNA sequencing.

Pyrosequencing technology is a bioluminometric DNA sequencing method that employs a cascade of four enzymes to deliver sequence signals. To date this technology has been limited to the sequencing of short stretches of DNA. As an improvement to this technique, we have introduced a bacterial group-specific, multiple sequencing primer approach that circumvents sequencing of less informative semi-conservative regions of the 16S rRNA gene. This new approach is suitable for challenging templates, improving sequence data quality, avoiding sequencing of non-specific amplification products, lessening sequencing time, and moreover, this strategy should open the way for many new applications in the future. The group-specific, multiple sequencing primers can be applied in the Sanger dideoxy sequencing method as well. In addition, we have improved the chemistry of the Pyrosequencing system enabling sequencing of longer stretches of DNA, which allows numerous new applications.

Bacteria↗

A 210-kb segment of tandem repeats and retroelements located between imprinted subdomains of mouse distal chromosome 7.

Mammalian genes subject to genomic imprinting often form clusters and are regulated by long-range mechanisms. The distal imprinted domain of mouse chromosome 7 is orthologous to the Beckwith-Wiedemann syndrome domain in human chromosome 11p15.5 and contains at least 13 imprinted genes. This domain consists of two subdomains, which are respectively regulated by an imprinting center. We here report the finished-quality sequence of a 0.6-Mb region encompassing the more centromeric subdomain. The sequence contains four imprinted genes (Ascl2/Mash2, Ins2, Igf2 and H19) and reveals previously unidentified CpG islands and tandem repeats, which may be features of imprinted genes. Most interestingly, a unique 210-kb segment consisting almost exclusively of tandem repeats and retroelements is identified. This segment, located between Th and Ins2, has features of heterochromatin-forming DNA and is highly methylated at CpG sites. The segment exhibits asynchronous replication on the parental chromosomes, a feature of the imprinted domains. We propose that this repeat segment could serve either as a boundary between the two subdomains or as a target for epigenetic chromatin modifications that regulate imprinting.

Animals↗

A comparative molecular analysis of developing mouse forelimbs and hindlimbs using serial analysis of gene expression (SAGE).

The analysis of differentially expressed genes is a powerful approach to elucidate the genetic mechanisms underlying the morphological and evolutionary diversity among serially homologous structures, both within the same organism (e.g., hand vs. foot) and between different species (e.g., hand vs. wing). In the developing embryo, limb-specific expression of Pitx1, Tbx4, and Tbx5 regulates the determination of limb identity. However, numerous lines of evidence, including the fact that these three genes encode transcription factors, indicate that additional genes are involved in the Pitx1-Tbx hierarchy. To examine the molecular distinctions coded for by these factors, and to identify novel genes involved in the determination of limb identity, we have used Serial Analysis of Gene Expression (SAGE) to generate comprehensive gene expression profiles from intact, developing mouse forelimbs and hindlimbs. To minimize the extraction of erroneous SAGE tags from low-quality sequence data, we used a new algorithm to extract tags from -analyzed sequence data and obtained 68,406 and 68,450 SAGE tags from forelimb and hindlimb SAGE libraries, respectively. We also developed an improved method for determining the identity of SAGE tags that increases the specificity of and provides additional information about the confidence of the tag-UniGene cluster match. The most differentially expressed gene between our SAGE libraries was Pitx1. The differential expression of Tbx4, Tbx5, and several limb-specific Hox genes was also detected; however, their abundances in the SAGE libraries were low. Because numerous other tags were differentially expressed at this low level, we performed a 'virtual' subtraction with 362,344 tags from six additional nonlimb SAGE libraries to further refine this set of candidate genes. This subtraction reduced the number of candidate genes by 74%, yet preserved the previously identified regulators of limb identity. This study presents the gene expression complexity of the developing limb and identifies candidate genes involved in the regulation of limb identity. We propose that our computational tools and the overall strategy used here are broadly applicable to other SAGE-based studies in a variety of organisms. [SAGE data are all available at GEO (http://www.ncbi.nlm.nih.gov/geo/) under accession nos. GSM55 and GSM56, which correspond to the forelimb and hindlimb raw SAGE data.]

Animals↗

An SNP resource for rice genetics and breeding based on subspecies indica and japonica genome alignments.

Dense coverage of the rice genome with polymorphic DNA markers is an invaluable tool for DNA marker-assisted breeding, positional cloning, and a wide range of evolutionary studies. We have aligned drafts of two rice subspecies, indica and japonica, and analyzed levels and patterns of genetic diversity. After filtering multiple copy and low quality sequence, 408,898 candidate DNA polymorphisms (SNPs/INDELs) were discerned between the two subspecies. These filters have the consequence that our data set includes only a subset of the available SNPs (in particular excluding large numbers of SNPs that may occur between repetitive DNA alleles) but increase the likelihood that this subset is useful: Direct sequencing suggests that 79.8% +/- 7.5% of the in silico SNPs are real. The SNP sample in our database is not randomly distributed across the genome. In fact, 566 rice genomic regions had unusually high (328 contigs/48.6 Mb/13.6% of genome) or low (237 contigs/64.7 Mb/18.1% of genome) polymorphism rates. Many SNP-poor regions were substantially longer than most SNP-rich regions, covering up to 4 Mb, and possibly reflecting introgression between the respective gene pools that may have occurred hundreds of years ago. Although 46.2% +/- 8.3% of the SNPs differentiate other pairs of japonica and indica genotypes, SNP rates in rice were not predictive of evolutionary rates for corresponding genes in another grass species, sorghum. The data set is freely available at http://www.plantgenome.uga.edu/snp.

Breeding↗

A comparison of rice chloroplast genomes.

Using high quality sequence reads extracted from our whole genome shotgun repository, we assembled two chloroplast genome sequences from two rice (Oryza sativa) varieties, one from 93-11 (a typical indica variety) and the other from PA64S (an indica-like variety with maternal origin of japonica), which are both parental varieties of the super-hybrid rice, LYP9. Based on the patterns of high sequence coverage, we partitioned chloroplast sequence variations into two classes, intravarietal and intersubspecific polymorphisms. Intravarietal polymorphisms refer to variations within 93-11 or PA64S. Intersubspecific polymorphisms were identified by comparing the major genotypes of the two subspecies represented by 93-11 and PA64S, respectively. Some of the minor genotypes occurring as intravarietal polymorphisms in one variety existed as major genotypes in the other subspecific variety, thus giving rise to intersubspecific polymorphisms. In our study, we found that the intersubspecific variations of 93-11 (indica) and PA64S (japonica) chloroplast genomes consisted of 72 single nucleotide polymorphisms and 27 insertions or deletions. The intersubspecific polymorphism rates between 93-11 and PA64S were 0.05% for single nucleotide polymorphisms and 0.02% for insertions or deletions, nearly 8 and 10 times lower than their respective nuclear genomes. Based on the total number of nucleotide substitutions between the two chloroplast genomes, we dated the divergence of indica and japonica chloroplast genomes as occurring approximately 86,000 to 200,000 years ago.

Base Sequence↗