Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Identification of peptides within a known protein sequence using COMSEQ analysis of data containing multiple sequences.

Modern methods of automated protein sequence analysis can provide high-quality data with which unambiguous amino-acid sequences can be determined, but analyses are more difficult when the sample is not pure. COMSEQ and auxillary programs were written to facilitate reconciliation of multiple amino-acid sequences potentially contained in noisy data with the known amino-acid sequence of the parent protein. The COMSEQ program prints a matrix in which the first vertical column represents the known amino-acid sequence of a selected protein. Each row of the matrix contains the sequencer yield corresponding to the amino acid in the first column, with each column corresponding to the sequencing reaction cycle. A diagonal which contains net increases of amino acids for each amino acid in the known sequence identifies a peptide potentially contained within the data. The number of matches for each diagonal over the entire known sequence are tabulated and presented as an aid to locating comparisons of greatest interest. The RNDSEQ program conducts multiple analyses using randomized versions of the known amino-acid sequence and tabulates the cumulative frequencies of potential sequence matches irrespective of the true known sequence. TRANSEQ is a utility program that translates edited sequence data from common databases into files that can be used by COMSEQ and RNDSEQ. The programs have been used successfully to identify two co-sequenced peptides from bovine serum albumin, an albumin peptide sequence in the presence of hemoglobin, and to identify two sequences of rat alpha-2u-globulin that differ in their amino termini.

Amino Acid Sequence↗

Signals for the selection of a splice site in pre-mRNA. Computer analysis of splice junction sequences and like sequences.

To evaluate the importance of the surrounding nucleotide sequence in the selection of a splice site for mRNA, we have carried out computer studies of eukaryotic protein genes whose entire nucleotide sequences were available. A splice site-like sequence that has a significant homology to the consensus splice junction sequences is frequently found within an intron and exon. It is found that the higher the homology of a candidate donor site sequence to the nine-nucleotide consensus sequence, the higher is its probability of being a donor site. For most of the donors, the stability of presumed base-pairing with U1-RNA is higher than that of donor-like sequences, if any, in the adjacent exon and intron. However, homology of a candidate acceptor sequence to the 15-nucleotide consensus is a poor criterion of an acceptor site. The presence of a sequence that could serve as a branch-point 18 to 37 nucleotides before an acceptor does not seem to be critical in distinguishing it from an acceptor-like sequence. For genes of human, rat, mouse and chicken, respectively, nucleotide frequencies around splice junctions of many genes have been calculated. They seem to be different at some positions around a donor site from species to species. The acceptors for these vertebrates have longer pyrimidine-rich regions than the previous consensus sequence. The newly derived nucleotide frequencies were used as the standard to calculate the weighted homology score of a candidate splice site sequence in a gene of the four species. This weighted homology score of the 40 to 60-nucleotide intron-exon sequence is a much better criterion of an acceptor. These results suggest that the most important signal in the selection of a splice resides in the surrounding nucleotide sequence. It is also suggested that the surrounding nucleotide sequence alone is not generally sufficient for the selection.

Animals↗

Nucleotide sequences of Caenorhabditis elegans core histone genes. Genes for different histone classes share common flanking sequence elements.

We have determined the nucleotide sequence of core histone genes and flanking regions from two of approximately 11 different genomic histone clusters of the nematode Caenorhabditis elegans. Four histone genes from one cluster (H3, H4, H2B, H2A) and two histone genes from another (H4 and H2A) were analyzed. The predicted amino acid sequences of the two H4 and H2A proteins from the two clusters are identical, whereas the nucleotide sequences of the genes have diverged 9% (H2A) and 12% (H4). Flanking sequences, which are mostly not similar, were compared to identify putative regulatory elements. A conserved sequence of 34 base-pairs is present 19 to 42 nucleotides 3' of the termination codon of all the genes. Within the conserved sequence is a 16-base dyad sequence homologous to the one typically found at the 3' end of histone genes from higher eukaryotes. The C. elegans core histone genes are organized as divergently transcribed pairs of H3-H4 and H2A-H2B and contain 5' conserved sequence elements in the shared spacer regions. One of the sequence elements, 5' CTCCNCCTNCCCACCNCANA 3', is located immediately upstream from the canonical TATA homology of each gene. Another sequence element, 5' CTGCGGGGACACATNT 3', is present in the spacer of each heterotypic pair. These two 5' conserved sequences are not present in the promoter region of histone genes from other organisms, where 5' conserved sequences are usually different for each histone class. They are also not found in non-histone genes of C. elegans. These putative regulatory sequences of C. elegans core histone genes are similar to the regulatory elements of both higher and lower eukaryotes. The coding regions of the genes and the 3' regulatory sequences are similar to those of higher eukaryotes, whereas the presence of common 5' sequence elements upstream from genes of different histone classes is similar to histone promoter elements in yeast.

Animals↗

Identification of novel transcribed sequences on human chromosome 22 by expressed sequence tag mapping.

To identify sequences on the human genome that are actually transcribed, we mapped expressed sequence tags (ESTs) of long cDNAs ranging from 4 kb to 7 kb along a 33.4-Mb sequence of human chromosome 22, the first human chromosome entirely sequenced. By the EST mapping of 30,683 long cDNAs in silico, 603 cDNA sequences were found to locate on chromosome 22 and classified into 169 clusters. Comparison of the genomic loci of these cDNA sequences with 679 genes already annotated on chromosome 22q revealed that 46 clusters represented newly identified transcribed sequences. To further characterize these sequences, we sequenced 12 cDNAs in their entirety out of 46 clusters. Of these 12 cDNAs, 6 were predicted to include a protein-coding region while the remaining 6 were unlikely to encode proteins. Interestingly, 3 out of the 12 cDNAs had the nucleotide sequences of the opposite strands of the genes previously annotated, which suggested that these genomic regions were transcribed bi-directionally. In addition to these newly identified 12 cDNAs, another 12 cDNAs were entirely sequenced since these cDNAs were likely to contain new information about the predicted protein-coding sequences previously annotated. In the cases of KIAA1670 and KIAA1672, these single cDNA sequences covered two separately annotated transcribed regions. For example, the sequence of a clone for KIAA1670 indicated that the CHKL and CPT1B genes were co-transcribed as a contiguous transcript without making both the protein-coding regions fused. In conclusion, the mapping of ESTs derived from long cDNAs followed by sequencing of the entire cDNAs provided indispensable information for the precise annotation of genes on the genome together with ESTs derived from short cDNAs.

Brain Chemistry↗

[MR-coronary angiography: comparison of SSFP and spoiled GRE sequence (bright blood technique) and a TSE sequence (black blood technique) in healthy volunteers].

PURPOSE: Comparison of a free breathing steady-state free precession (SSFP), a spoiled gradient-echo (GRE) and a turbo spin-echo sequence (TSE) for imaging of the coronary arteries (MRCA) in healthy volunteers. MATERIALS AND METHODS: Twenty-two healthy volunteers were imaged with a standard clinical scanner (1.5 T, Intera, Philips), with the right coronary system imaged in 11 and the left coronary system in the other 11 volunteers. Images were obtained with a 3D-SSFP (balanced TFE, TR 6.2 ms, TE 3.1 ms, alpha 65 degrees ), a 3D-GRE (TFE, TR 7.2 ms, TE 2.2 ms, alpha 30 degrees ) and a 2D-TSE (Dual-IR, TR 2RR, TE 25 ms) sequence. The in plane resolution was 0.7 x 0.8 mm for both the SSFP and GRE sequence with an effective slice thickness of 1.5 mm. For the TSE sequence, an in-plane resolution of 0.7 x 0.9 mm and a slice thickness of 3.0 mm were used. All investigations were performed using prospective navigator gating and slice-following technique. The signal-to-noise ratio (SNR) and contrast-to-noise ratio (CNR) for the blood pool to myocardium and blood pool to epicardial fat were calculated. Image quality and measurement artifacts were assessed for all sequences by 5 independent investigators using a 4- and 5-point grading scale. RESULTS: CNR was significantly higher for the GRE sequence compared with the SSFP sequence and TSE sequence (mean 20.8 +/- 4.8 vs. 14.6 +/- 5.0 and 10.1 +/- 3.7 for blood pool to myocardium; mean 27.5 +/- 6.3 vs. 16.4 +/- 5.4 and 18.1 +/- 5.7 for blood pool to fat). The SNR revealed no significant differences between the SSFP and GRE sequences. The SSFP and the TSE sequences showed significantly more artefacts than the spoiled GRE sequence. Image quality was graded slightly higher for the GRE than for the SSFP sequence for the right coronary system, while there was no substantial difference in the left coronary system (median 2.1 +/- 0.6 and 2.5 +/- 0.6 vs. 2.5 +/- 0.8 and 2.6 +/- 0.7 for the right and left coronary system). In comparison, image quality was lower with the TSE sequence (median 2.9 +/- 0.5 for the right coronary system with p < 0.05 vs. GRE sequence and 3.0 +/- 0.3 for the left coronary system). CONCLUSION: For the scan parameters chosen in this study, the GRE-sequence represents the most robust technique for imaging of the coronary arteries. Currently, the TSE sequence is no alternative.

Adult↗

Effect of sequence length on the execution of familiar keying sequences: lasting segmentation and preparation?

The author assessed the mechanisms underlying skilled production of keying sequences in the discrete sequence-production task by examining the effect of sequence length on mean element execution rate (i.e., the rate effect). To that end, participants (N = 9) practiced fixed movement sequences consisting of 2, 4, and 6 key presses for a total of 588 trials per sequence. In the subsequent test phase, the sequences were executed with and without a verbal short-term memory task in both simple and choice reaction time (RT) paradigms. The rate effect was obtained in the discrete sequence-production task-including the typical quadratic increase in sequence execution time (SET, which excludes RT) with sequence length. The rate effect resulted primarily from 6-key sequences that included 1 or 2 relatively slow elements at individually different serial positions. Slowing of the depression of the 2nd response key (R2) in the 2-key sequence reduced the rate effect in the memory task condition, and faster execution of the 1st few elements in each sequence amplified the rate effect in simple RT. Last, the time to respond to random cues increased with position, suggesting that the mechanisms that underlie the rate effect in new sequences and in familiar sequences are different. The data were in line with the notion that coding of longer keying sequences involves motor chunks for the individual sequence segments and information on how those motor chunks are to be concatenated.

Adult↗

Coronavirus transcription mediated by sequences flanking the transcription consensus sequence.

In our studies of murine coronavirus transcription, we continue to use defective interfering (DI) RNAs of mouse hepatitis virus (MHV) in which we insert a transcription consensus sequence in order to mimic subgenomic RNA synthesis from the nondefective genome. Using our subgenomic DI system, we have studied the effects of sequences flanking the MHV transcription consensus sequence on subgenomic RNA transcription. We obtained the following results. (i) Insertion of a 12-nucleotide-long sequence including the UCUAAAC transcription consensus sequence at different locations of the DI RNA resulted in different efficiencies of subgenomic DI RNA synthesis. (ii) Differences in the amount of subgenomic DI RNA were defined by the sequences that flanked the 12-nucleotide-long sequence and were not affected by the location of the 12-nucleotide-long sequence on the DI RNA. (iii) Naturally occurring flanking sequences of intergenic sequences at gene 6-7, but not at genes 1-2 and 2-3, contained a transcription suppressive element(s). (iv) Each of three naturally occurring flanking sequences of an MHV genomic cryptic transcription consensus sequence from MHV gene 1 also contained a transcription suppressive element(s). These data showed that sequences flanking the transcription consensus sequence affected MHV transcription.

Animals↗

Dog Y chromosomal DNA sequence: identification, sequencing and SNP discovery.

BACKGROUND: Population genetic studies of dogs have so far mainly been based on analysis of mitochondrial DNA, describing only the history of female dogs. To get a picture of the male history, as well as a second independent marker, there is a need for studies of biallelic Y-chromosome polymorphisms. However, there are no biallelic polymorphisms reported, and only 3200 bp of non-repetitive dog Y-chromosome sequence deposited in GenBank, necessitating the identification of dog Y chromosome sequence and the search for polymorphisms therein. The genome has been only partially sequenced for one male dog, disallowing mapping of the sequence into specific chromosomes. However, by comparing the male genome sequence to the complete female dog genome sequence, candidate Y-chromosome sequence may be identified by exclusion. RESULTS: The male dog genome sequence was analysed by Blast search against the human genome to identify sequences with a best match to the human Y chromosome and to the female dog genome to identify those absent in the female genome. Candidate sequences were then tested for male specificity by PCR of five male and five female dogs. 32 sequences from the male genome, with a total length of 24 kbp, were identified as male specific, based on a match to the human Y chromosome, absence in the female dog genome and male specific PCR results. 14437 bp were then sequenced for 10 male dogs originating from Europe, Southwest Asia, Siberia, East Asia, Africa and America. Nine haplotypes were found, which were defined by 14 substitutions. The genetic distance between the haplotypes indicates that they originate from at least five wolf haplotypes. There was no obvious trend in the geographic distribution of the haplotypes. CONCLUSION: We have identified 24159 bp of dog Y-chromosome sequence to be used for population genetic studies. We sequenced 14437 bp in a worldwide collection of dogs, identifying 14 SNPs for future SNP analyses, and giving a first description of the dog Y-chromosome phylogeny.

Animals↗

Average values of a dissimilarity measure not requiring sequence alignment are twice the averages of conventional mismatch counts requiring sequence alignment for a computer-generated model system.

Three measures of sequence dissimilarity have been compared on a computer-generated model system in which substitutions in random sequences were made at randomly selected sites and the replacement character was chosen at random from the set of characters different from the original occupant of the site. The three measures were the conventional mismatch count between aligned sequences (AMC = m) and two measures not requiring prior sequence alignment. The latter two measures were the squared Euclidean distance between vectors of counts of t-tuples (t = 1-6) of characters in the two sequences (multiplet distribution distances or MDD = d) and counts of characters not covered by word structures of statistically significant length common to the two sequences (common long words or CLW = SIB, SIS, or SAB). Average MDD distances were found to be two times average mismatch counts in the simulated sequences for all values of t from 1 to 6 and all degrees of substitution from one per sequence to so many as to produce, effectively, random sequences. This simple relation held independently of sequence length and of sequence composition. The relation was confirmed by exact results on small model systems and by formal asymptotic results in the limit of so few substitutions that no double hits occur and in the limit of two random sequences. The coefficient of variation for MDD distances was greater than that for mismatch counts for singlets but both measures approached the same low value for sextets. Needleman-Wunsch alignment produced incorrect mismatch counts at higher degrees of substitution. The model satisfied the conditions for the derivation of the Jukes-Cantor asymptotic adjustment, but its application produced increasingly bad results with increasing degrees of substitution in accord with earlier results on model and natural sequences. This fact was a consequence of the increase with increasing degrees of substitution of the sensitivity of the adjustment to error in the observations. Average CLW distances for a variety of common word structures were more or less parallel to MDD distances for appropriately long t-tuples. These results on model systems supported the validity of the two dissimilarity measures not requiring sequence alignment that was found in earlier work on natural sequences (Blaisdell 1989).

Base Sequence↗

An evaluation of the significance of amino acid sequence homologies in human histocompatibility antigens (HLA-A and HLA-B) with immunoglobulins and other proteins, using relatively short sequences.

A computer search was carried out for homologies between HLA-A and HLA-B antigen sequences and the sequences of constant and variable regions of immunoglobulins and of all other sequenced proteins. Searches were made both with relatively short peptide sequences from the HLA antigens and with those longer peptide sequences which were available in 1978. Significant homology of HLA antigen sequences to immunoglobulin constant region sequences was found in two cases: (1) a short decapeptide sequence which includes the fourth cysteine residue of HLA-B7 and (2) an 89-amino-acid residue (Ac-2) C-terminal fragment of the papainsolubilized HLA-B7 molecule. The difficulty of establishing statistically significant sequence homology with relatively short peptide sequences is emphasized by computer-based comparisons of the decapeptide sequence with randomly generated peptide sequences. It is concluded that statistically significant homology with short sequences can be assured only when extraordinarily high degrees of homology are present and additional constraints are included in the matches, for example, matches at relatively rare amino acid residues such as Cys, His and Trp. The homology of the 89-amino-acid residue sequence to constant region sequences of immunoglobulins is as great as or greater than that of beta 2-microglobulin. These findings and the unique domain structure involving a disulphide loop of comparable size strongly favour a common evolutionary origin for this region of HLA-A and -B, beta 2-microglobulin and immunoglobulin constant regions.

Amino Acid Sequence↗

Discovery of diverse anellovirus sequences in Thai human sequencing data.

UNLABELLED: Anelloviruses are part of the normal human viral flora. Although their diversity in humans has been investigated in many countries, and despite their initial detection in Thailand in 1999, knowledge of Thai anelloviruses remains very limited. This study analyzed 1,175 whole-genome sequencing data sets from Thai individuals to mine for potential anellovirus sequences. Our analyses detected anellovirus sequences in 149 data sets (12.68%), uncovering 434 partial anellovirus sequences and 77 complete genome sequences, characterized by the presence of terminal redundancy, complete orf1, and the conserved untranslated region upstream of the orf1 gene. Sequence analyses indicated that these viruses belong to seven genera, including Alphatorquevirus, Betatorquevirus, Gammatorquevirus, Hetorquevirus, Lamedtorquevirus, Samektorquevirus, and Yodtorquevirus. Notably, Hetorquevirus, Lamedtorquevirus, Samektorquevirus, and Yodtorquevirus had not previously been reported in Thailand. Phylogenetic analysis of ORF1 protein sequences showed that Thai anelloviruses form multiple phylogenetic clusters with non-Thai anelloviruses, indicating frequent cross-country transmission and multiple origins of the virus in Thailand. Furthermore, sequence similarity network analysis identified 33 potentially novel anellovirus species in our data set. Our findings greatly expand the knowledge of anellovirus diversity in Thailand and demonstrate the potential of human whole-genome sequencing data as a valuable resource for viral discovery. Lastly, we highlight and discuss some challenges with the use of the current pairwise sequence similarity-based classification scheme, in particular, how gaps can influence similarity calculation and potentially lead to inconsistencies with a phylogenetic-based classification scheme. IMPORTANCE: Anelloviruses are widespread in humans, yet their diversity remains poorly characterized in many regions, including Thailand. Here, we demonstrate that human sequencing data sets, originally generated without the intention for virome research, can be effectively mined for anellovirus sequences, including complete genomes. Our findings reveal a substantial number of previously unreported anelloviruses in Thailand, significantly expanding the known diversity of the virus. We also highlight potential limitations of the current anellovirus species classification scheme, which is based on pairwise orf1 sequence similarity analysis with a hard threshold cutoff at 69%. Our results reveal that the current scheme can sometimes yield taxonomic groupings that are inconsistent with phylogenetic relationships, particularly when significant alignment gaps are present. Overall, our results show that existing human sequencing data can be effectively repurposed for virus discovery research and suggest the need for more robust and phylogenetically informed classification frameworks as viral sequence databases continue to expand.

Humans↗

Sequence length and error analysis of Sequenase and automated Taq cycle sequencing methods.

We have examined DNA sequence error as a function of length using both a manual method of performing reactions with Sequenase and an automated Taq cycle sequencing method. DNA fragments from both methods were separated and analyzed on a sequencer. To determine the sequence of a cosmid insert (35.3 kb), 379 sequences were obtained from a manual Sequenase method, and 354 sequences were obtained from a Taq cycle sequencing method as performed on an automated robotic workstation and sequenced on an automated fluorescent sequencer. A highly redundant consensus of these sequences was obtained and aligned with the individual sequences to determine sequence error over the length of each sequence. The results of this study indicate that error is about 1% per position over the first 350 nucleotides, but increases thereafter to about 17% at 500 nucleotides. This pattern of accuracy was nearly equivalent for manual Sequenase methods and automated Taq cycle sequencing methods. The potential of these methods in large-scale DNA sequencing projects is discussed.

Cloning, Molecular↗

Intermediate sequences increase the detection of homology between sequences.

Two homologous sequences, which have diverged beyond the point where their homology can be recognised by a simple direct comparison, can be related through a third sequence that is suitably intermediate between the two. High scores, for a sequence match between the first and third sequences and between the second and the third sequences, imply that the first and second sequences are related even though their own match score is low. We have tested the usefulness of this idea using a database that contains the sequences of 971 protein domains whose structures are known and whose residue identities with each other are some 40% or less (PDB40D). On the basis of sequence and structural information, 2143 pairs of these sequences are known to have an evolutionary relationship. FASTA, in an all-against-all comparison of the sequences in the database, detected 320 (15%) of these relationships as well as three false positive (i.e. 1% error rate). Using intermediate sequences found by FASTA matches of PDB40D sequences to those in the large non-redundant OWL database we could detect 550 evolutionary relationships with an error rate of 1%. This means the intermediate sequence procedure increases the ability to recognise the evolutionary relationships amongst the PDB40D sequences by 70%.

Amino Acid Sequence↗

Use of an automated sequencer to determine the sequence specificity of DNA damage.

An automated sequencer was used to determine the sequence specificity of DNA damage caused by hedamycin in the plasmid pUC19 using a linear amplification/Taq DNA polymerase method. Previously, manual DNA sequencers have been in widespread use to investigate the sequence specificity of a DNA damaging agent. Manual DNA sequencers are restricted in the length of DNA sequence that can be read at base pair resolution for densitometry. An automated sequencer can greatly expand on the length of analysable DNA sequence. An additional important capability of the automated sequencer, is the ability to quantitate the intensity of damage at each base pair site. Thus we have used the automated sequencer to elucidate the sequence specificity of DNA damage for 300 bp. We have carried out an extended analysis of the sequence specificity of hedamycin DNA damage and found that the sequence 5'-cGt-3', tGt and cGg are preferentially damaged. The sequence specificity of cisplatin was also investigated.

Alkylating Agents↗

Typing of Candida glabrata in clinical isolates by comparative sequence analysis of the cytochrome c oxidase subunit 2 gene distinguishes two clusters of strains associated with geographical sequence polymorphisms.

We tested whether comparative sequence analysis of the mitochondrion-encoded cytochrome c oxidase subunit 2 gene (COX2) could be used to distinguish intraspecific variants of Candida glabrata. Mitochondrial genes are suitable for investigation of close phylogenetic relationships because they evolve much faster than nuclear genes, which in general exhibit very limited intraspecific variation. For this survey we used 11 clinical isolates of C. glabrata from three different geographical locations in Brazil, 10 isolates from one location in the United States, 1 American Type Culture Collection strain as an internal control, and the published sequence of strain CBS 138. The complete coding region of COX2 was amplified from total cellular DNA, and both strands were sequenced twice for each strain. These sequences were aligned with published sequences from other fungi, and the numbers of substitutions and phylogenetic relationships were determined. Typing of these strains was done by using 17 substitutions, with 8 being nonsynonymous and 9 being synonymous. Also, cDNAs made from purified mitochondrial polyadenylated RNA were sequenced to confirm that our sequences correspond to the expressed copies and not nuclear pseudogenes and that a frameshift mutation exists in the 3' end of the coding region (position 673) relative to the Saccharomyces cerevisiae sequence and the previously published C. glabrata sequence. We estimated the average evolutionary rate of COX2 to be 11.4% sequence divergence/10(8) years and that phylogenetic relationships of yeasts based on these sequences are consistent with rRNA sequence data. Our analysis of COX2 sequences enables typing of C. glabrata strains based on 13 haplotypes and suggests that positions 51 and 519 indicate a geographical polymorphism that discriminates strains isolated in the United States and strains isolated in Brazil. This provides for the first time a means of typing of Candida strains that cause infections by use of direct sequence comparisons and the associated divergence estimates.

Bacterial Proteins↗

[The full sequence of intron 51 of dystrophin gene and its characteristic of sequence].

OBJECTIVE: To finish the work of sequencing the full sequence of intron 51 of dystrophin gene and understand its characteristic of sequence. METHODS: The whole intron 51 was sequenced by primer walking. The sequencing results were analyzed by repeat sequences, matrix attachment region (MAR) and topoisomerase II cleavage sites. The residue sequences, after removal of the repetitive sequences, were subjected to the analysis of CpG islands, promoter, open reading frame (ORF) and unidentified low copy repeat sequence. RESULTS: The acquired intron 51 sequence was composed of 38725 bp. Repetitive sequences constituted 37.53% of total intron sequence. The overall G+C content of intron 51 was 36.34%. There are four potential MARs in intron 51. Three of them are clustered in the 12 kb region near exon 51. Numerous ORFs were found on both strands, but no homologues proteins were found in Genbank CDS transcriptional peptide, PDB, SwissProt, PIR and PRF databases. CONCLUSION: The expansion of intron 7 over the last 120 million years was mainly the result of L1 insertion into intron 7, and not all of repetitive sequences are associated with chromosomal rearrangement. No sequence of functional significance was found in intron 51. The results suggest that the cluster of MARs may be associated with the instability of intron 51.

Base Sequence↗