Search PubMedSearch

SEARCH · Search PubMed

Results for “Terminal Repeat Sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

A model for replication of the ends of linear chromosomes.

Linear chromosomes possessing internal repeats of their terminal sequences can form intramolecular crossed-strand exchanges that allow replication of the chromosome ends. Evidence is discussed that such a mechanism may be utilized during replication of herpes simplex virus DNA and during replication of macronuclear DNA from the hypotrichous ciliate Oxytricha.

Animals

Chromosome-level genome assembly of the ornamental plant Alcea rosea.

Alcea rosea, a member of the Malvaceae family, is celebrated for its rich floral palette and global horticultural significance. Here, we present a high-quality reference genome for A. rosea, achieving a genome assembly size of 1.01 Gbp, with a Contig N50 length of 36.61 Mbp. The genome sequence was successfully mapped to 21 chromosomes, and the scaffold N50 length reached 52.57 Mbp, with a scaffold genome completeness of 99.6%. A total of 565.84 Mbp (comprising 56% of the genome) of repetitive sequences were identified, with transposable elements being predominant, particularly long terminal repeat (LTR) elements, which accounted for 48.44% of the genome. 51,436 genes were annotated. Among these predicted genes, the average gene length and coding sequence (CDS) length were 2739.92 bp and 1242.54 bp, respectively.

Genome, Plant

Sequence arrangement in herpes simplex virus type 1 DNA: identification of terminal fragments in restriction endonuclease digests and evidence for inversions in redundant and unique sequences.

It has been proposed by Sheldrick and Berthelot (1974) that the terminal sequences of herpes simplex virus type 1 (HSV-1) DNA are repeated in an internal inverted form and that the inverted redundant sequences delimit and separate two unique sequences, S and L. In this study the sequence arrangement in HSV-1 DNA has been investigated with restriction endonuclease cleavage, end-labeling studies, and molecular hybridization experiments. The terminal fragments in digests with restriction endonucleases Hind III, Hpa-1, EcoRI and Bum were identified and shown to be consistent with the Sheldrick and Berthelot model. Inverted fragments which contain unique sequences as well as redundant sequences, and which the model predicts, were identified by DNA-DNA hybridization studies. Further cleavage of Bum fragments with Hpa-1 also revealed inversions of the terminal sequences that contained unique sequences. The results obtained showed that the unique sequences S and L are relatively inverted in different DNA molecules in the population, resulting in the presence of four related genomes with rearranged sequences in apparently equal amounts. The redundant sequences bounding S do not share complete sequence homology with those bounding L, but hybridization studies are presented which show that the terminal 0.3% of the genome is repeated in every redundant sequence.

Base Sequence

Structure and organization of the two tRNATyr gene clusters on the E. coli chromosome.

The structure and organization of the gene clusters coding for the two tyrosine-accepting tRNA species (tRNA1Tyr and tRNA2Tyr) on the E. coli chromosome have been determined. The mature structural sequences of the two tRNATyr genes, located on opposite sides of the E. coli chromosome, differ by only 2 bp, but sequences surrounding these portions of the genes are very different. The genes coding for tRNA1Tyr (tyrT) comprise two mature structural sequences separated by a 200 bp "intergenic spacer." It is known that in transducing phage, the region adjoining the CCA end of the second mature structural sequence comprises a 178 bp repeated sequence which contains an in vitro, rho-dependent transcriptional termination site. We find that these potentially genetically unstable repeated sequences are present in the E. coli chromosome with the same organization as that determined from transducing phage analyses. The gene that codes for tRNA2Tyr (tyrU) is present in a single copy and is tightly clustered with three other tRNA genes. One of these genes (to be called thrU) encodes a previously undescribed tRNA (to be called tRNA4Thr). The organization of this cluster on the E. coli chromosome is tRNA4Thr--8 bp--tRNA2Tyr--115 bp--tRNA2Gly--6 bp--tRNA3Thr. The importance of correlating structural analyses derived from specialized transducing phage with those determined for the chromosome itself is demonstrated by results which show that out of four independently isolated tRNATyr transducing phage, two carrying the tRNA1Tyr genes [phi80psu3+,- (Cambridge) and phi80sus2psu3+ (Kyoto)] and two carrying the tRNA2Tyr gene (lambdarifd 18 and lambdah80dglyTsu+36), only the first phage from each group has the same gene organization as that found in the E. coli chromosome.

Base Sequence

A chromosome-level genome assembly and annotation of Cercis chuniana (Fabaceae).

The genus Cercis L., at the base of the subfamily Cercidoideae of Fabaceae, is known for its ecological adaptability and significant medicinal, ornamental, and economic value. However, the lack of a high-quality genome hinders the understanding of the evolution of Cercis and Fabaceae. In this study, we present a chromosome-level genome of Cercis chuniana by combining Illumina short reads, PacBio HiFi long reads, and Hi-C data. The final genome size is 355.53 Mb, consisting of 12 contigs with a N50 of 42.34 Mb. Notably, 344.24 Mb, corresponding to 96.82% of the genome, was anchored to seven chromosomes. The assembly comprises 24.83% repetitive sequences, including 19.32% long terminal repeats. Additionally, a total of 33,837 protein-coding genes were predicted in the genome, with 32,709 (96.67%) genes successfully annotated. The high-quality genome assembly of C. chuniana not only bridges the existing gap in genomic data and offers important resources for molecular studies of this species, but also provides essential insights for future studies on speciation, functional and comparative genomics within the Fabaceae family.

Genome, Plant

Designer TALEs enable discovery of cell death-inducer genes.

Transcription activator-like effectors (TALEs) in plant-pathogenic Xanthomonas bacteria activate expression of plant genes and support infection or cause a resistance response. PthA4AT is a TALE with a particularly short DNA-binding domain harboring only 7.5 repeats which triggers cell death in Nicotiana benthamiana; however, the genetic basis for this remains unknown. To identify possible target genes of PthA4AT that mediate cell death in N. benthamiana, we exploited the modularity of TALEs to stepwise enhance their specificity and reduce potential target sites. Substitutions of individual repeats suggested that PthA4AT-dependent cell death is sequence specific. Stepwise addition of repeats to the C-terminal or N-terminal end of the repeat region narrowed the sequence requirements in promoters of target genes. Transcriptome profiling and in silico target prediction allowed the isolation of two cell death inducer genes, which encode a patatin-like protein and a bifunctional monodehydroascorbate reductase/carbonic anhydrase protein. These two proteins are not linked to known TALE-dependent resistance genes. Our results show that the aberrant expression of different endogenous plant genes can cause a cell death reaction, which supports the hypothesis that TALE-dependent executor resistance genes can originate from various plant processes. Our strategy further demonstrates the use of TALEs to scan genomes for genes triggering cell death and other relevant phenotypes.

Cell Death

A Chromosome-Level Genome Assembly of the Potato Leafhopper Empoasca fabae (Hemiptera: Cicadellidae).

The potato leafhopper, Empoasca fabae (Harris, 1841), is a highly polyphagous, migratory insect pest of eastern North America that feeds on more than 200 herbaceous and woody plant species, causing substantial losses to forage and field crops. Despite its agricultural and ecological importance, no genome has been available for this species. Here, we present the first chromosome-level genome assembly of E. fabae, generated from Oxford Nanopore long reads, Illumina short reads, and Omni-C proximity-ligation data. The final assembly spans 908 Mb across 132 scaffolds, with 99.8% of the assembly captured in ten chromosome-length scaffolds (nine autosomes and an X chromosome) with a scaffold N50 of 96.2 Mb. The assembly is highly complete, recovering 92.9% of conserved hemipteran single-copy orthologs from protein annotations, and is composed of 47.6% repetitive sequence, dominated by long terminal repeat retrotransposons and unclassified elements. Read-depth comparison between male and female individuals supports assignment of a single sex-linked chromosome, consistent with an XO sex determination system. BRAKER3 gene annotation predicted 31,406 protein-coding genes after retaining the longest isoform per locus. Comparative genome analysis of the two closest related Typhlocybinae species with genomes available, Matsumurasca onukii and Hebata decipiens, revealed extensive chromosome-scale collinearity while defining a shared core gene repertoire. This reference genome provides a foundation for comparative and population genomic studies and for investigating genetic traits in this economically important crop pest species.

Animals

Characterization of long and short repetitive sequences in the sea urchin genome.

Long and short repetitive sequences were purified from the DNA of Paracentrotus lividus under conditions designed to optimize the yield of complete, end to end sequences. Double-stranded long repeat DNA prepared in this manner ranged in length from approximately 3000 to 15 000 nucleotide pairs with average sizes of approximately 6000 base pairs. In the electron microscope, long repeat DNA was observed to possess continuous sequences that often appeared to be terminated by one or more loops and/or fold backs. Long repeat DNA sequences, resheared to 300 base pairs, were found to have an average melting point identical to that for sheared native DNA. Thus, the reassociated duplexes of long repetitive DNA seem to possess very few mismatched base pairs. Reassociation kinetic analyses indicate that the majority of the long repeat sequences are reiterated only 4--7 times per haploid amount of DNA. Melt-reassociation analyses of short repetitive DNA, at several criteria, support the previously held concept that these sequences belong the sets or families of sequences which are inexact copies of one another. Our studies also support hypotheses suggesting that short repetitive sequences belong to families which may have arisen via distinct salttatory events. The relationships between long and short repetitive DNA sequences are considered with respect to widely held concepts of their sequence organization, evolution, and possible functions within eucaryotic genomes. A model for the possible organization of short repeats within long repetitive DNA sequences is also presented.

Animals

Physical maps for Herpes simplex virus type 1 DNA for restriction endonucleases Hind III, Hpa-1, and X. bad.

It has been proposed that the genome of herpes simplex virus type 1 (HSV-1) consists of two internal unique sequences, S and L, bounded by two sets of redundant sequences (P. Sheldrick and N. Berthelot, 1974). In this arrangement, terminal sequences (TRs and TRl) are repeated in an internal inverted form (IRs and IRl) and delimit S and L. Furthermore, a body of evidence has accumulated that suggests that S and L themselves are inverted, giving rise to four related forms of the HSV genome. In this study the ordering of restruction endonuclease fragments of HSV-1 DNA for physical maps has been studied using molecular hybridization techniques and the cleavage of isolated restriction endonuclease fragments with further restriction endonucleases. Physical maps for the fragments produced by Hind III, Hpa-1, and X. bad have been constructed for the four related forms of the HSV-1 genome. TRs and IRs were found to be between 3.5 x 10(6) and 4.5 x 10(6) daltons, TRl and IRl about 6 x 10(6) daltons, S about 8 x 10(6) to 9 x 10(6) daltons, and L about 6.8 x 10(6) daltons.

Base Sequence

The piRNA pathway mediates transcriptional silencing of LTR retrotransposons in ovaries and somatic tissues of Aedes mosquitoes.

The PIWI-interacting RNA (piRNA) pathway preserves genomic integrity by suppressing transposable elements in animal germlines. Despite its well-established function in the animal germline, piRNAs and PIWI proteins are expressed in somatic tissues across arthropod species, and their functions outside the gonads remain poorly understood. Aedes albopictus mosquitoes express four PIWI genes, Piwi4, Piwi5, Piwi6, and Ago3, in both gonadal and somatic tissues. Here, we generated Piwi6 knockout (KO) Ae. albopictus cell lines and observed a substantial upregulation of long terminal repeat retrotransposons, including a full-length endogenous retrovirus that we named Aedes albopictus Endogenous Retrovirus-1 (AalERV1). Nascent RNA sequencing and Cleavage Under Targets and Tagmentation (CUT&Tag) analyses revealed that Piwi6 silences AalERV1 transcriptionally by guiding the deposition of the repressive H3K9me3 histone mark. Consistently, Piwi6 localized to both the cytoplasm and nucleus, with sequences in the intrinsically disordered region guiding nuclear translocation. Reintroduction of full-length GFP-Piwi6, but not a mutant GFP-Piwi6 defective in nuclear localization, rescued AalERV1 repression in Piwi6 KO cells. Importantly, Piwi6-mediated control of AalERV1 was recapitulated in vivo as Piwi6 knockdown increased AalERV1 expression in both ovaries and somatic tissues of Ae. albopictus mosquitoes. These results establish Aedes mosquitoes as a model to study nuclear PIWI functions and suggest that somatic piRNA-mediated transposon silencing is evolutionarily conserved across arthropod species.

Animals

Transposition of elements of the 412, copia and 297 dispersed repeated gene families in Drosophila.

The stability of elements of three different dispersed repeated gene families in the genome of Drosophila tissue culture cells has been examined. Different amounts of sequences homologous to elements of 412, copia and 297 dispersed repeated gene families are found in the genomes of D. melanogaster embryonic and tissue culture cells. In general the amount of these sequences is increased in the cell lines. The additional sequences homologous to 412, copia and 297 occur as intact elements and are dispersed to new sites in the cell culture genome. It appears that these elements can insert at many alternative sites. We also describe a DNA sequence arrangement found in the D. melanogaster embryo genome which appears to result from a transposition of an element of the copia dispersed repeated gene family into a new chromosomal site. The mechanism of insertion of this copia element is precise to within 90 bp and may involve a region of weak sequence homology between the site of insertion and the direct terminal repeats of the copia element.

Animals

Ribosomal RNA genes of Saccharomyces cerevisiae. II. Physical map and nucleotide sequence of the 5 S ribosomal RNA gene and adjacent intergenic regions.

A DNA fragment containing the structural gene for the 5 S ribosomal RNA and intergenic regions before and after the 35 S ribosomal RNA precursor gene of Saccharomyces cerevisiae has been amplified in a bacterial plasmid and physically mapped by restriction endonuclease cleavage and hybridization to purified yeast 5 S ribosomal RNA. The nucleotide sequence of the DNA fragments carrying the 5 S ribosomal RNA gene and adjacent regions has been determined. The sequence unambiguously identifies the 5 S ribosomal RNA gene, determines its polarity within the ribosomal DNA repeating unit, and reveals the structure of its promoter and termination regions. Partial DNA sequence of the regions near the beginning and end of the 35 S ribosomal RNA gene has also been determined as a preliminary step in establishing the structure of promoter and termination regions for the 35 S ribosomal RNA gene.

Base Sequence

Hide and seek: de novo identification in sugar beet reveals impact of non-autonomous LTR retrotransposons.

Plant genomes are filled with retrotransposons and their derivatives, constantly undergoing sequence diversification and structural rearrangement. Among them, short, non-autonomous retrotransposons lack full coding capacity and often form subfamilies. As a result, non-autonomous retrotransposons are incompletely identified in most to all genome assemblies.Here, we capitalize on our comprehensive understanding of the transposable element (TE) landscape in sugar beet (Beta vulgaris) to assess the extent of the blind spot for non-autonomous long terminal repeat (LTR) retrotransposons. This use case serves to answer if all of these sequences are derivatives of easier-to-identify full-length elements or if there is more variability that is currently overlooked.For this we applied a semi-automated structural discovery workflow followed by in-depth manual verification to characterize non-autonomous LTR retrotransposons in sugar beet. We retrieve more than 100 non-autonomous LTR retrotransposon families that lack complete autonomous coding capacity, including canonical terminal-repeat retrotransposons in miniature (TRIMs), elongated non-coding derivatives and families retaining fragmented coding remnants. The identified families span a broad range, including elements exceeding 15,000 bp in length and display evidence for reshuffling and modular evolution. Only a subset of families could be confidently linked to autonomous retrotransposons, showing sequence diversification within the non-autonomous LTR retrotransposon fraction beyond the autonomous genomic templates.We highlight that a large fraction of non-autonomous LTR retrotransposons is incompletely recovered with the current TE identification workflows, even if the output is well-curated and condensed into TE libraries and suggest procedures to remedy this gap. This study gives a genome-wide view into the non-autonomous LTR retrotransposon landscape of a single plant genome and highlights the importance of structure-based approaches for their identification and classification.

LTR retrotransposons

Visualization of an inverted terminal repetition in vaccinia virus DNA.

An inverted terminal repetition was observed in DNA molecules extracted from vaccinia virus. The repeated sequence was visualized by (i) nicking the hairpin loops present of the ends of vaccinia virus DNA, (ii) separating the strands of DNA by alkali denaturation, (iii) allowing the single strands to self-anneal, and (iv) examining the DNA with an electron microscope. Single-stranded circular molecules, each of which contained a duplex projection (3.54 +/- 0.12 micron) representing the terminal repetition, readily formed. Similar size projections were also seen in heteroduplex structures formed by crosshybridization of the separated strands of the two terminal HindIII restriction fragments. Based on contour length measurements and the electrophoretic mobility of the isolated inverted terminal repetition, a molecular weight of approximately 6.9 X 10(6), equivalent to about 10,500 nucleotide base pairs, was estimated. Evidence was obtained from DNA-RNA hybridization studies that the terminal repetition is transcribed.

Base Sequence

Nucleotide sequence of the genes III, VI and I of bacteriophage M13.

A DNA region of 2750 base pairs encompassing the genes III, VI and I of bacteriophage M13 has been sequenced by the Maxam-Gilbert procedure. By establishing the nucleotide changes introduced by several amber mutations, the coding region and the regulatory signals of each gene have been deduced. The genes appear to span 1275 base pairs (gene III; mol.wt. 44,748) 339 base pairs (gene VI; mol.wt. 12,264) and 1047 base pairs (gene I; mol.wt. 39,500). Their separating non-codogenic regions are extremely short, namely two and one base pair, respectively. The C-terminal end of gene I, however, intrudes 23 nucleotides into gene IV. From the nucleotide sequence it appears that the minor capsid protein of the phage, which is encoded by gene III, is synthesized in a precursor form containing 18 extra amino acids at its N-terminal end. Furthermore, in this capsid protein two clusters of a fourfold repeat of the sequence Glu-Gly-Gly-Gly-Ser are apparent. Gene VI appears to code for a small, extremely hydrophobic polypeptide. Its total hydrophobic amino acids content of 51% suggests that this protein can only function in the host cell membrane.

Amino Acid Sequence

Suppression and reversal of allergic encephalomyelitis in guinea pigs with a non-encephalitogenic analogue of the tryptophan region of the myelin basic protein.

The administration of synthetic peptide S42 leads to suppression and reversal of experimental allergic encephalomyelitis (EAE) induced in guinea pigs by myelin basic protein. Peptide S42 contains a linear sequence of 21 amino acid residues, H-Phe-Ser-Trp-Gln-Lys-Phe-Ser-Trp-Gln-Lys-Phe-Ser-Trp-Gln-Lys-Phe-Ser-Trp-Gln-Lys-Gly-OH, made up of four repeating unit sequences of H-Phe-Ser-Trp-Gln-Lys-OH in addition to a C-terminal glycine. Injected at relatively high doses, peptide S42 is non-encephalitogenic. It induces delayed-type hypersensitivity which is not followed by EAE, and elicits delayed-type hypersensitivity responses in peptide S42, encephalitogenic trytophan peptide, or BP-challenged animals for either of the three antigens. The repeating unit sequence of peptide S42 is analogous to the encephalitogenic tryptophan region of the BP molecules . The sequence homology is responsible for cellular recognition of this antigen by the skin test assay and suggests in vivo interaction between peptide S42 and EAE-inducing cells leading to suppression and reversal of disease.

Animals

Megamimivirus double-stranded DNA linear genomes flanked by highly diverse terminal inverted repeats.

UNLABELLED: Giant viruses have fundamentally expanded our understanding of virology by challenging the conventional boundaries of both virion size and genome complexity. However, the scarcity of isolates has left many of their unique biological features unexplored. Here, we report the isolation and characterization of four new giant virus species belonging to the subfamily Megamimivirinae, sampled from distinct environments across China. Among these, Megavirus daqingense is the first giant virus isolated from an oil reservoir; it exhibits virion stability under high salinity, chloroform exposure, and elevated temperatures, suggesting fitness adaptations to subsurface conditions. Using a hybrid sequencing approach that integrates short- and long-read technologies, we assembled complete linear genomes for all four isolates, each flanked by long terminal inverted repeats (TIRs). Comparative genomic and synteny analyses identified 29 distinct TIRs from 46 megamimivirus genomes. Gene content within these TIRs was highly diverse, with no orthologous proteins conserved across all repeats. Furthermore, TIR genes experienced weaker purifying selection than those in non-TIR regions (i.e., the genomic regions excluding the TIRs), consistent with their role as drivers of genome plasticity. Notably, we discovered for the first time that identical tRNA genes are shared between TIRs and non-TIR regions of eukaryotic viruses. Collectively, our work provides insights into the structural and evolutionary complexity of megamimiviruses, revealing TIRs as reservoirs of genetic diversity and hotspots for gene transfer, thereby playing a pivotal role in shaping the dynamic architecture of giant virus genomes. IMPORTANCE: Terminal inverted repeats (TIRs) are critical structural elements at the termini of linear genomes essential for fundamental processes such as recombination, replication, and integration across diverse organisms. However, the inherent limitations of short-read sequencing technologies have left the complete structure, diversity, and evolutionary significance of long TIRs in giant viruses unexplored. In this study, we leverage hybrid sequencing and comparative genomic analyses to unveil the complexity of TIRs across the subfamily Megamimivirinae. We demonstrate that TIRs are dynamic genomic hotspots characterized by remarkable gene diversity and unexpected conservation of specific tRNA genes. These findings establish TIRs as key drivers of genome plasticity, serving as hotspots for horizontal gene transfer and genetic innovation. By resolving the long-hidden terminal structures of megamimivirus genomes, this work provides a foundational framework for understanding how TIRs shape the evolution of giant viruses and, more broadly, advances our understanding of genome architecture in large DNA viruses.

Megavirus

Proviruses of avian sarcoma virus are terminally redundant, co-extensive with unintegrated linear DNA and integrated at many sites.

We have analyzed the DNA from 15 clones of avian sarcoma virus (ASV)-transformed rat cells with restriction endonucleases and molecular hybridization techniques to determine the location and structure of proviral DNA. All twenty units of proviral DNA identified in these 15 clones appear to be inserted at different sites in host DNA. In each of the ten cases that could be sufficiently well mapped, entirely different regions of cellular DNA were involved. Thus ASV DNA can be accommodated at many positions in cellular DNA, but the existence of preferred sites has not been excluded. Six of the 15 clones carry only one normal provirus, two contain two normal proviruses, and seven harbor either one or two proviruses that appear anomalous in physical mapping tests. Both ends of at least 18 proviruses, however, were found to contain sequences specific to both the 3' and 5' termini of viral RNA. The organization of these terminally redundant sequences appeared identical to that of the 300 base pair (bp) repeats found at the ends of unintegrated linear DNA (Shank et al., 1978). Proviral DNA is therefore co-extensive, or nearly co-extensive, with unintegrated linear DNA and has a structure we denote as CELL DNA-3'5'----------3'5'-CELL DNA. Three of the four anomalous proviruses which were fully analyzed were deletion mutants lacking 25--65% of the genetic content of ASV; the fourth provirus had a novel site for cleavage by Eco RI but was otherwise normal. Tests for the biological competence of proviral DNA, based upon rescue of transforming virus after fusion with chicken cells, were generally consistent with the physical mapping studies.

Avian Sarcoma Viruses