Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

A new method that simultaneously aligns and reconstructs ancestral sequences for any number of homologous sequences, when the phylogeny is given.

Among the fundamental problems in molecular evolution and in the analysis of homologous sequences are alignment, phylogeny reconstruction, and the reconstruction of ancestral sequences. This paper presents a fast, combined solution to these problems. The new algorithm gives an approximation to the minimal history in terms of a distance function on sequences. The distance function on sequences is a minimal weighted path length constructed from substitutions and insertions-deletions of segments of any length. Substitutions are weighted with an arbitrary metric on the set of nucleotides or amino acids, and indels are weighted with a gap penalty function of the form gk = a + (bxk), where k is the length of the indel and a and b are two positive numbers. A novel feature is the introduction of the concept of sequence graphs and a generalization of the traditional dynamic sequence comparison algorithm to the comparison of sequence graphs. Sequence graphs ease several computational problems. They are used to represent large sets of sequences that can then be compared simultaneously. Furthermore, they allow the handling of multiple, equally good, alignments, where previous methods were forced to make arbitrary choices. A program written in C implemented this method; it was tested first on 22 5S RNA sequences.

Algorithms↗

Murine mammary tumor virus pol-related sequences in human DNA: characterization and sequence comparison with the complete murine mammary tumor virus pol gene.

Sequences in the human genome with homology to the murine mammary tumor virus (MMTV) pol gene were isolated from a human phage library. Ten clones with extensive pol homology were shown to define five separate loci. These loci share common sequences immediately adjacent to the pol-like segments and, in addition, contain a related repeat element which bounds this region. This organization is suggestive of a proviral structure. We estimate that the human genome contains 30 to 40 copies of these pol-related sequences. The pol region of one of the cloned segments (HM16) and the complete MMTV pol gene were sequenced and compared. The nucleotide homology between these pol sequences is 52% and is concentrated in the terminal regions. The MMTV pol gene contains a single long open reading frame encoding 899 amino acids and is demarcated from the partially overlapping putative gag gene by termination codons and a shift in translational reading frame. The pol sequence of HM16 is multiply terminated but does contain open reading frames which encode 370, 105, and 112 amino acid residues in separate reading frames. We deduced a composite pol protein sequence for HM16 by aligning it to the MMTV pol gene and then compared these sequences with other retroviral pol protein sequences. Conserved sequences occur in both the amino and carboxyl regions which lie within the polymerase and endonuclease domains of pol, respectively.

Amino Acid Sequence↗

Nucleotide sequence of Rhizobium meliloti insertion sequence ISRm1: homology to IS2 from Escherichia coli and IS426 from Agrobacterium tumefaciens.

Nucleotide sequencing of Rhizobium meliloti insertion sequence ISRm1 showed that it is 1319 nucleotides long and includes 32/31 nucleotide terminal inverted repeats. Analysis of five different insertion sites using sequencing primers complementary to sequences within the left and right ends demonstrated that ISRm1 generates five bp direct repeats at the sites of insertion. Although ISRm1 has shown a target preference for certain short regions (hot spots), there was no apparent similarity in the DNA sequences near the insertion sites. On one strand ISRm1 contains two contiguous open reading frames (ORFs) spanning most of its length. ISRm1 was found to have over 50% sequence homology to insertion sequences IS2 from Escherichia coli and IS426 from Agrobacterium tumefaciens. Their sizes, the sequences of their inverted repeats, and the characteristics of their insertion sites are also comparable, indicating that ISRm1, IS2 and IS426 belong to a class of related insertion sequences. Comparison of the proteins potentially encoded by these insertion sequences showed that the two ORFs found in ISRm1 are also present in IS2 and IS426, suggesting that they may be functional genes.

Agrobacterium tumefaciens↗

High performance DNA sequencing, and the detection of mutations and polymorphisms, on the Clipper sequencer.

The Visible Genetics Clipper sequencer is a new platform for automated DNA sequencing which employs disposable MicroCel cassettes and 50 microm thick polyacrylamide gels. Two DNA ladders can be analyzed simultaneously in each of 16 lanes on a gel, after labeling with far-red absorbing dyes such as Cy5 and Cy5.5. This allows a simultaneous bidirectional sequencing of four templates. We have evaluated the Clipper sequencer, by cycle-sequencing of an M13 single-stranded DNA standard, and by coupled amplification and sequencing (CLIP) of reverse-transcribed human immunodeficiency virus (HIV-1) RNA standards and clinical patient samples. (i) Limitations of instrument. We have examined basic instrument parameters such as detector stability, background, digital sampling rate, and gain. With proper usage, the optical and electronic subsystems of the Clipper sequencer do not limit the data collection or sequence-determination processes. (ii) Limitations of gel performance. We have also examined the physics of DNA band separation on 50 microm thick MicroCel gels. We routinely obtain well-resolved sequence which can be base-called with 98.5% accuracy to position approximately 450 on an 11 cm gel, and to position approximately 900 on a 25 cm gel. Resolution on 5 and 11 cm gels ultimately is limited by a sharp decrease in spacing between adjacent bands, in the biased reptation separation regime. Fick's (thermal) diffusion appears to be of minor importance on 6 cm or 11 cm gels, but becomes an additional resolution-limiting factor on 25 cm gels. (iii) Limitations of enzymology. Template quality, primer nesting, choice of DNA polymerase, and choice between dye primers and dye terminators are key determinants of the ability to detect mutations and polymorphisms on the Clipper sequencer, as on other DNA sequencers. When CLIP is used with dye-labeled primers and a DNA polymerase of the F667Y, delta(5'--> 3' exo) class, we can routinely detect single-nucleotide mutations and polymorphisms over the 0.35-0.65 heterozygosity range. We present an example of detecting therapeutically relevant mutations in a clinical HIV-1 RNA isolate.

Bacteriophage M13↗

The evolution of proteins from random amino acid sequences: II. Evidence from the statistical distributions of the lengths of modern protein sequences.

This paper continues an examination of the hypothesis that modern proteins evolved from random heteropeptide sequences. In support of the hypothesis, White and Jacobs (1993, J Mol Evol 36:79-95) have shown that any sequence chosen randomly from a large collection of nonhomologous proteins has a 90% or better chance of having a lengthwise distribution of amino acids that is indistinguishable from the random expectation regardless of amino acid type. The goal of the present study was to investigate the possibility that the random-origin hypothesis could explain the lengths of modern protein sequences without invoking specific mechanisms such as gene duplication or exon splicing. The sets of sequences examined were taken from the 1989 PIR database and consisted of 1,792 "super-family" proteins selected to have little sequence identity, 623 E. coli sequences, and 398 human sequences. The length distributions of the proteins could be described with high significance by either of two closely related probability density functions: The gamma distribution with parameter 2 or the distribution for the sum of two exponential random independent variables. A simple theory for the distributions was developed which assumes that (1) protoprotein sequences had exponentially distributed random independent lengths, (2) the length dependence of protein stability determined which of these protoproteins could fold into compact primitive proteins and thereby attain the potential for biochemical activity, (3) the useful protein sequences were preserved by the primitive genome, and (4) the resulting distribution of sequence lengths is reflected by modern proteins. The theory successfully predicts the two observed distributions which can be distinguished by the functional form of the dependence of protein stability on length. The theory leads to three interesting conclusions. First, it predicts that a tetra-nucleotide was the signal for primitive translation termination. This prediction is entirely consistent with the observations of Brown et al. (1990a,b, Nucleic Acids Res 18:2079-2086 and 18: 6339-6345) which show that tetra-nucleotides (stop codon plus following nucleotide) are the actual signals for termination of translation in both prokaryotes and eukaryotes. Second, the strong dependence of statistical length distributions on sequence-termination signaling codes implies that the evolution of stop codons and translation-termination processes was as important as gene splicing in early evolution. Third, because the theory is based upon a simple no-exon stochastic model, it provides a plausible alternative to a limited universe of exons from which all proteins evolved by gene duplication and exon splicing (Dorit et al. 1990, Science 250:1377-1382).

Amino Acid Sequence↗

Selection of DNA sequences from interval 6 of the human Y chromosome with homology to a Y chromosomal fertility gene sequence of Drosophila hydei.

An experimental approach towards the molecular analysis of the male fertility function, located in interval 6 of the human Y chromosome, is presented. This approach is not based on the knowledge of any gene product but on the assumption that the functional DNA structure of male fertility genes, evolutionary conserved with their position on the Y chromosome, may contain an evolutionary conserved frame structure or at least conserved sequence elements. We tested this hypothesis by using dhMiF1, a fertility gene sequence of the Y chromosome of Drosophila hydei, as a screening probe on a pool of cloned human Y-DNA sequences. We were able to select 10 human Y-DNA sequences of which 7 could be mapped to Y interval 6 (the pY6H sequence family). Since the only fertility gene of the human Y chromosome is mapped to the same Y interval, our working hypothesis seems to be strongly supported. Most interesting in this respect is the isolation of the Y-specific repetitive pY6H65 sequence. The pY6H65 locus extends to a length of at least 300 kb in Y interval 6 and has a locus-specific repetitive sequence organization, reminiscent of the functional DNA structure of Y chromosomal fertility genes of Drosophila. We identified the simple sequence family (CA)n as one sequence element conserved between the Drosophila dhMiFi fertility gene sequence and the homologous human Y-DNA sequences.

Animals↗

Nucleotide sequence of Xenopus borealis oocyte 5S DNA: comparison of sequences that flank several related eucaryotic genes.

Genomic Xenopus borealis oocyte-specific 5S DNA (Xbo) contains clusters of 5S rRNA genes. The number of genes varies among clusters, and the distance between genes within a cluster is about 80 nucleotides. The spacer DNA between gene clusters is AT-rich and heterogeneous in length due in part to variable numbers of a tandemly repeated 21 nucleotide sequence. A cloned fragment of Xbo 5S DNA (Xbo1) containing three 5S rRNA genes has been sequenced. The sequences of Xbo1 genes 1 and 2 are very similar to the dominant 5S RNA sequence, whereas 15 of the 120 residues in the third gene are different. The sequence of gene 3 is as different from the dominant gene sequence as the X. laevis pseudogene is from the 5S RNA gene. Sequence analysis of genomic DNA shows that gene 3 is an abundant component of the multigene family. All three genes are transcribed when added to an extract of X. laevis oocyte nuclei, and a fragment of Xbo1 lacking the AT-rich spacer DNA and the 5' end of the first gene supports transcription of genes 2 and 3 in this in vitro system. Thus the 80 nucleotides preceding each 5S gene are sufficient for promoter function. Nucleic acid sequences preceding several eucaryotic genes that are transcribed by RNA polymerase III were analyzed and the following common features were found: a purine-rich region; at least one direct repeat; the absence of dyad symmetry; transcription beginning with a purine; a pyrimidine residue immediately preceding the first nucleotide of the gene; and the oligonucleotides AAAAG, AGAAG and GAC, located approximately 15, 25 and 35 nucleotides, respectively, before the start of transcription. The 10 base pair (bp) spacing between the homologous oligonucleotides is that expected for a recognition signal on one face of a DNA double helix. The extensive sequence differences between most of the spacers that precedes these genes make the three conserved oligonucleotides more striking. Parts of the 5' flanking regions of the three Xbo1 gene (-12 to -40), which include the conserved oligonucleotides, are identical. In contrast, 7 of the first 11 nucleotides that precede the third 5S RNA gene in Xbo1 differ from those that precede the first gene. The sequences following the X. borealis oocyte and somatic 5S genes are identical in 12 of the first 14 residues and contain two or more T clusters, as does the corresponding region of X. laevis oocyte 5S DNA. The 3' sequences of the Xenopus 5S rRNA genes and several other eucaryotic genes contain features in common with procaryotic transcription termination sites. The 3' end of the gene is GC-rich and contains a dyad symmetry. Termination occurs in an AT-rich region containing one or more T clusters on the noncoding strand.

Animals↗

3DCoffee: combining protein sequences and structures within multiple sequence alignments.

Most bioinformatics analyses require the assembly of a multiple sequence alignment. It has long been suspected that structural information can help to improve the quality of these alignments, yet the effect of combining sequences and structures has not been evaluated systematically. We developed 3DCoffee, a novel method for combining protein sequences and structures in order to generate high-quality multiple sequence alignments. 3DCoffee is based on TCoffee version 2.00, and uses a mixture of pairwise sequence alignments and pairwise structure comparison methods to generate multiple sequence alignments. We benchmarked 3DCoffee using a subset of HOMSTRAD, the collection of reference structural alignments. We found that combining TCoffee with the threading program Fugue makes it possible to improve the accuracy of our HOMSTRAD dataset by four percentage points when using one structure only per dataset. Using two structures yields an improvement of ten percentage points. The measures carried out on HOM39, a HOMSTRAD subset composed of distantly related sequences, show a linear correlation between multiple sequence alignment accuracy and the ratio of number of provided structure to total number of sequences. Our results suggest that in the case of distantly related sequences, a single structure may not be enough for computing an accurate multiple sequence alignment.

Protein Conformation↗

Multiple group-specific sequencing primers for reliable and rapid DNA sequencing.

Pyrosequencing technology is a bioluminometric DNA sequencing method that employs a cascade of four enzymes to deliver sequence signals. To date this technology has been limited to the sequencing of short stretches of DNA. As an improvement to this technique, we have introduced a bacterial group-specific, multiple sequencing primer approach that circumvents sequencing of less informative semi-conservative regions of the 16S rRNA gene. This new approach is suitable for challenging templates, improving sequence data quality, avoiding sequencing of non-specific amplification products, lessening sequencing time, and moreover, this strategy should open the way for many new applications in the future. The group-specific, multiple sequencing primers can be applied in the Sanger dideoxy sequencing method as well. In addition, we have improved the chemistry of the Pyrosequencing system enabling sequencing of longer stretches of DNA, which allows numerous new applications.

Bacteria↗

The 'shortmer' approach to nucleic acid sequence analysis. I: Computer simulation of sequencing projects to find economical primer sets.

In principle it is most economical to sequence large DNA fragments consecutively ('primer walking'), provided there is an immediate supply of sequencing primers. To solve the problem of primer supply we previously suggested generating a bank of short oligonucleotide primers ('shortmers'). In every sequencing reaction shortmers would have to be selected from this bank that are suitable to hybridize adjacently on the sequencing template. After their ligation the shortmers would form a long, and hence more specific, primer in the subsequent sequencing reaction. In the present study a computer simulation of large sequencing projects revealed a reduced set of approximately 12,000 selected octanucleotides (out of all 65,536) retaining maximum priming flexibility and minimum redundant information on the simulated sequence analyses. Establishing routine protocols for nucleic acid sequencing following the shortmer approach will abolish the tightest bottleneck of the consecutive sequencing route (primer supply) and hence may render this general scheme more attractive than the shotgun sequencing scheme. A twofold (or more) speed-up of genome sequencing projects by the shortmer approach may be assumed.

Algorithms↗

Long double-stranded sequences (dsRNA-B) of nuclear pre-mRNA consist of a few highly abundant classes of sequences: evidence from DNA cloning experiments.

DNA preparations from about hundred randomly selected clones containing mouse DNA fragments were screened for the existence of sequences complementary to long double-stranded regions of pre-mRNA able to snap back after melting (dsRNA-B). Many clones containing such sequences were found. The cloned sequences can be subdivided into three groups: (1) those complementary to about a half (at least to 30-40%) of the total dsRNA, designated as sequences B1; (2) those complementary to a part of sequence B1; and (3) sequences complementary to about a quarter (at least to 15%) of the total dsRNA referred to as sequence B2. The size of DNA sequence complementary to dsRNA is about 400 base pairs. Melting experiments with hybrids show that the members of B1 family are very similar if not identical, while the divergence among B2 sequences is higher, but still the number of substitutions does not exceed 9% of bases. Thus, the major part of dsRNA-B consists of a small number of highly abundant sequences as was suggested earlier on the basis of renaturation kinetics /1-3/. Sequences B1 and B2 are represented by many copies in the mouse genome and in pre-mRNA, and many of them probably do not form hairpin-like structures.

Animals↗

The 10 kb Drosophila virilis 28S rDNA intervening sequence is flanked by a direct repeat of 14 base pairs of coding sequence.

Most repeat units of rDNA in Drosophila virilis are interrupted in the 28S rRNA coding region by an intervening sequence about 10 kb in length; uninterrupted repeats have a length of about 11 kb. We have sequenced the coding/intervening sequence junctions and flanking regions in two independent clones of interrupted rDNA, and the corresponding 28S rRNA coding region in a clone of uninterrupted rDNA. The intervening sequence is terminated at both ends by a direct repeat of a fourteen nucleotide sequence that is present once in the corresponding region of an intact gene. This is a phenomenon associated with transposable elements in other eukaryotes and in prokaryotes, and the Drosophila rDNA intervening sequence is discussed in this context. We have compared more than 200 nucleotides of the D. virilis 28S rRNA gene with sequences of homologous regions of rDNA in Tetrahymena pigmentosa (Wild and Sommer, 1980) and Xenopus laevis (Gourse and Gerbi, 1980): There is 93% sequence homology among the diverse species, so that the rDNA region in question (about two-thirds of the way into the 28S rRNA coding sequence) has been very highly conserved in eukaryote evolution. The intervening sequence in T. pigmentosa is at a site 79 nucleotides upstream from the insertion site of the Drosophila intervening sequence.

Animals↗

From mapping to sequencing, post-sequencing and beyond.

The Rice Genome Research Program (RGP) in Japan has been collaborating with the international community in elucidating a complete high-quality sequence of the rice genome. As the pioneer in large-scale analysis of the rice genome, the RGP has successfully established the fundamental tools for genome research such as a genetic map, a yeast artificial chromosome (YAC)-based physical map, a transcript map and a phage P1 artificial chromosome (PAC)/bacterial artificial chromosome (BAC) sequence-ready physical map, which serve as common resources for genome sequencing. Among the 12 rice chromosomes, the RGP is in charge of sequencing six chromosomes covering 52% of the 390 Mb total length of the genome. The contribution of the RGP to the realization of decoding the rice genome sequence with high accuracy and deciphering the genetic information in the genome will have a great impact in understanding the biology of the rice plant that provides a major food source for almost half of the world's population. A high-quality draft sequence (phase 2) was completed in December 2002. Since then, much of the finished quality sequence (phase 3) has become available in public databases. With the completion of sequencing in December 2004, it is expected that the genome sequence would facilitate innovative research in functional and applied genomics. A map-based genome sequence is indispensable for further improvement of current rice varieties and for development of novel varieties carrying agronomically important traits such as high yield potential and tolerance to both biotic and abiotic stresses. In addition to genome sequencing, various related projects have been initiated to generate valuable resources, which could serve as indispensable tools in clarifying the structure and function of the rice genome. These resources have been made available to the scientific community through the Rice Genome Resource Center (RGRC) of the National Institute of Agrobiological Sciences (NIAS) to enable rapid progress in research that will lead to thorough understanding of the rice plant. As the next trend in rice genome research will focus on determining the function of about 40,000-50,000 genes predicted in the genome as well as applying various genomics tools in rice breeding, an unlimited access to rice DNA and seed stocks will provide a broad community of scientists with the necessary materials for formulating new concepts, developing innovative research and making new scientific discoveries in rice genomics.

Centromere↗

DNA sequencing with [alpha-33P]-labeled ddNTP terminators: a new approach to DNA sequencing with Thermo Sequenase DNA polymerase.

A new approach to DNA sequencing is described. The method is based on the use of [alpha-33P]-labeled dideoxyribonucleoside triphosphate terminators and Thermo Sequenase DNA polymerase in cycle sequencing. Thermo Sequenase DNA polymerase incorporates ddNTPs as efficiently as dNTPs, allowing the use of low concentrations of these nucleotides in DNA sequencing. Because only the properly terminated chains are labeled and visualized on autoradiography of the sequencing gels, the sequence results are free of background. The intensity of DNA bands generated are remarkably uniform, which makes reading of DNA sequences easy. By staggered loading of the sequencing gel (at 2-3 hour intervals), it is possible to sequence DNA at least 450 to 500 nucleotides. Exposure time for autoradiography with [alpha-33P] labels is much shorter than with [35S] and does not substantially compromise autoradiographic resolution. Data can be obtained after only 12 hours of exposure of an X-ray film. Moreover, cycle sequencing requires very small amounts of single- or double-stranded template. Consequently, it is even possible to generate sequence data from a single bacterial colony. The details of the protocol are presented in a stepwise manner, and some important parameters to be considered for sequencing with this method are discussed.

Codon, Terminator↗

Characterization of the repetitive sequences in a 200-kb region around the rice waxy locus: diversity of transposable elements and presence of veiled repetitive sequences.

Repetitive genomic sequences might have various structural features and properties distinct from those of the known transposable elements (TE). Here, the content and properties of the repetitive sequences present in a 200-kb region around the rice waxy locus were analyzed using the available rice genomic database. In our previous Southern blotting analysis, 70% of the segments in this region showed smeared patterns, but according to the present database analysis, the proportion of repetitive sequences in this region was only 15%. The repetitive segments in this 200-kb region comprised 75 repetitive sequences that we classified into 46 subfamilies: 21 subfamilies were known TEs or repetitive sequences and 25 subfamilies consisted of newly identified TEs or novel types of repetitive sequences. The region contains no long terminal repeat (LTR) retrotransposable elements, but miniature inverted repeat transposable elements (MITEs) constituted a major class among the elements identified. These MITEs showed remarkable structural divergence: 12 elements were found to be new members of known MITE superfamilies, while five elements had novel terminal structures, and did not belong to any known TE families. Interestingly, about 10% of the repetitive sequences, including virus-like sequences did not have any of the usual characteristics of TEs, suggesting that a certain proportion of repetitive sequences that might not share the transpositional mechanisms of known elements are dispersed in the compact rice genome.

Base Sequence↗

The oli1 gene and flanking sequences in mitochondrial DNA of Saccharomyces cerevisiae: the complete nucleotide sequence of a 1.35 kilobase petite mitochondrial DNA genome covering the oli1 gene.

As part of our genetic and molecular analysis of mutants of Saccharomyces cerevisiae affected in the oli1 gene (coding for mitochondrial ATPase subunit 9) we have determined the complete nucleotide sequence of the mtDNA genome of a petite (23-3) carrying this gene. Petite 23-3 (1,355 base pairs) retains a continuous segment of the relevant wild-type (J69-1B) mtDNA genome extending 983 nucleotides upstream, and 126 nucleotides downstream, of the 231 nucleotide oli1 coding region. There is a 15-nucleotide excision sequence in petite 23-3 mtDNA which occurs as a direct repeat in the wild-type mtDNA sequence flanking the unique petite mtDNA segment (interestingly, this excision sequence in petite 23-3 carries a single base substitution relative to the parental wild-type sequence). The putative replication origin of petite 23-3 is considered to be in its single G,C rich cluster, which differs in just one nucleotide from the standard oriS sequence. The DNA sequences in the intergenic regions flanking the oli1 gene of strain J69-1B (and its derivatives) have been systematically compared to those of the corresponding regions of mtDNA in strains derived from the D273-10B parent (sequences from the laboratory of A. Tzagoloff). The nature and distribution of the sequence divergences (base substitutions, base deletions or insertions, and more extensive rearrangements) are considered in the context of functions associated with mitochondrial gene expression which are ascribed to specialized sequences in the intergenic regions of the yeast mitochondrial genome.

Base Sequence↗

Sequence and evolution of related bovine and caprine satellite DNAs. Identification of a short DNA sequence potentially involved in satellite DNA amplification.

The satellite II DNAs of the domestic ox Bos taurus and sheep Ovis aries have been sequenced, and that of the domestic goat Capra hircus partially sequenced. All three are related, and consist of repeat units of about 700 base-pairs. There is no evidence of internal repetition within these repeat units. When matched for maximum homology, the goat and sheep sequences show 83% homology, whereas the ox and sheep sequences share only 70% homology. Factors contributing to the uncertainty of the exact homology between these sequences are discussed, but the results are nevertheless consistent with their progenitor sequence being present in the common ancestor of cattle and sheep. Goat satellite II DNA is shown to contain another, unrelated, tandemly repeated sequence, which is composed of 22 base-pair repeat units. Both this sequence and a region of ox satellite II share good homology with the 11 base-pair progenitor sequence of ox 1.706 g/cm3 satellite DNA. It is suggested that this shared sequence could play a role in bovine satellite DNA amplification.

Animals↗

N sequences, P nucleotides and short sequence homologies at junctional sites in VH to VHDJH and VHDJH to JH joining.

Junctional sequences in VH to VHDJH and VHDJH to JH joining occurring in Abelson virus-transformed immature B cell lines were PCR-amplified and sequenced. In VH to VHDJH joining, 24 (23%) out of 105 junctions examined here had P nucleotides and/or N sequences, and out of the remaining 81 junctions without P nucleotides and N sequences, 57 (70%) had short sequence homologies of one or two bases (A, C, G, CA or AG) and three had long sequence homologies at the junctional sites. In VHDJH to JH joining, 38 (43%) out of 89 junctions examined here had P nucleotides and/or N sequences, and out of the remaining 51 junctions without P nucleotides and N sequences, 47 (92%) had short sequence homologies of one or two bases (C, T, G, TG or GG) at the junctional sites. These results indicate that short sequence homologies play an important role for end joining in VH to VHDJH and VHDJH to JH joining.

Animals↗