Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

The primary sequence and the subunit structure of mouse alpha-2-macroglobulin, deduced from protein sequencing of the isolated subunits and from molecular cloning of the cDNA.

Mouse plasma alpha-2-macroglobulin (m alpha 2M) was isolated and the N-terminal amino-acid sequences determined after separation of the 165-kDa and 35-kDa subunits. These sequences were compared to the protein sequence predicted by the cDNA, which was cloned from a mouse liver library and sequenced. From these data it is evident that both subunits are encoded by one mRNA of approximately 5 kb expressed predominantly in liver. The smaller subunit, with the N-terminal sequence DLSSSDLT, comprises the C-terminal 257 residues of m alpha 2M and is derived from a single-chain precursor probably by proteolytic processing at an arginine residue in the sequence PTRDLSS. Analysis of the predicted protein further showed all the salient features of a proteinase inhibitor of the macroglobulin family: a bait region that deviates from all known sequences in this family, a very conserved internal thiolester site and conserved cysteine residues and putative N-glycosylation sites. The synthesis of m alpha 2M in adult liver was demonstrated by Northern blotting and in fetal liver by in-situ hybridization. Transient transfection of COS cells with the cDNA under control of a viral promoter demonstrated the secretion and partial processing of m alpha 2M in the culture medium. In plasma the level of m alpha 2M was found to be stable as expected for the murine counterpart of human plasma alpha-2-macroglobulin. The possibilities of using the mouse as a genetic model to study this proteinase inhibitor in vivo are discussed.

Amino Acid Sequence↗

Variations of two repetitive DNA sequences in several Triticeae genomes revealed by polymerase chain reaction and sequencing.

Genomes of Triticeae were analyzed using PCR with synthesized primers that were based on two published repetitive DNA sequences, pLeUCD2 (pLe2) and 1-E6hcII-1 (L02368),which were originally isolated from Thinopyrum elongatum. The various genomes produced a 2240 bp PCR product having high homology with the repetitive DNA pLe2. The PCR fragments produced from different genomes differed mainly in amplification quantity and in base composition at 89 variable sites. On the other hand, amplification products from the primer set for L02368 were of different sizes and nucleotide sequences. These results show that the two repetitive DNA sequences have different evolutionary significance. ple2 is present in all genomes tested, although differences in copy number and nucleotide sequence are notable. L02368 is more genome specific, i.e., fewer genomes possess this family of repetitive sequences. It was concluded that the repetitive sequence pLe2 family is an ancient one that existed in progenitor genome prior to divergence of annual and perennial genomes. In contrast, sequences similar to L02368 have only evolved following genome divergence.

Base Sequence↗

Functional dissection of the Tol2 transposable element identified the minimal cis-sequence and a highly repetitive sequence in the subterminal region essential for transposition.

The Tol2 element is a naturally occurring active transposable element found in vertebrate genomes. The Tol2 transposon system has been shown to be active from fish to mammals and considered to be a useful gene transfer vector in vertebrates. However, cis-sequences essential for transposition have not been characterized. Here we report the characterization of the minimal cis-sequence of the Tol2 element. We constructed Tol2 vectors containing various lengths of DNA from both the left (5') and the right (3') ends and tested their transpositional activities both by the transient excision assay using zebrafish embryos and by analyzing chromosomal transposition in the zebrafish germ lineage. We demonstrated that Tol2 vectors with 200 bp from the left end and 150 bp from the right end were capable of transposition without reducing the transpositional efficiency and found that these sequences, including the terminal inverted repeats (TIRs) and the subterminal regions, are sufficient and required for transposition. The left and right ends were not interchangeable. The Tol2 vector carrying an insert of >11 kb could transpose, but a certain length of spacer, <276 but >18 bp, between the left and right ends was necessary for excision. Furthermore, we found that a 5-bp sequence, 5'-(A/G)AGTA-3', is repeated 33 times in the essential subterminal region. Mutations in the repeat sequence at 13 different sites in the subterminal region, as well as mutations in TIRs, severely reduced the excision activity, indicating that they play important roles in transposition. The identification of the minimal cis-sequence of the Tol2 element and the construction of mini-Tol2 vectors will facilitate development of useful transposon tools in vertebrates. Also, our study established a basis for further biochemical and molecular biological studies for understanding roles of the repetitive sequence in the subterminal region in transposition.

Animals↗

Isolation of cDNA including a reverse transcriptase-like sequence transcribed from the long interspersed repetitive DNA sequence of rat.

The mammalian genome congains long interspersed repetitive sequences, but the role of these repetitive sequence is not clear. A cDNA clone has been isolated that contains part of the L1 sequences from a cDNA library of rat liver. The DNA sequence analysis showed the homology of cDNA to several reverse transcriptases. The homology between the amino acid sequences predicted from L1 consensus sequences and reverse transcriptases has been reported previously. However, this is the first isolation of a cDNA clone containing a reverse transcriptase-like sequence.

Amino Acid Sequence↗

A high frequency of length polymorphisms in repeated sequences adjacent to Alu sequences.

We describe a new class of DNA length polymorphism that is due to a variation in the number of tandem repeats associated with Alu sequences (Alu sequence-related polymorphisms). The polymerase chain reaction was used to selectively amplify a (TTA)n repeat identified in the 3-hydroxy-3-methylglutaryl coenzyme A (HMG CoA) reductase gene from genomic DNA of 41 human subjects, and the size of the amplified products was determined by gel electrophoresis. Seven alleles were found that differed in size by integrals of three nucleotides. The allele frequencies ranged from 1.5% to 52%, and the overall heterozygosity index was 62%. The polymorphic TTA repeat was located adjacent to a repetitive sequence of the Alu family. A homology search of human genomic DNA sequences for the trinucleotide TTA (at least five members in length) revealed tandem repeats in six other genes. Three of the six (TTA)n repeats were located adjacent to Alu sequences, and two of the three (in the genes for beta-tubulin and interleukin-1 alpha) were found to be polymorphic in length. Tandemly repetitive sequences found in association with Alu sequences may be frequent sites of length polymorphism that can be used as genetic markers for gene mapping or linkage analysis.

Base Sequence↗

[Plan for finding homologies in nucleotide sequence databases using preliminarily calculated sequence samples].

A scheme of fast similarity search of nucleotide sequences is suggested based on sequence imaging, which results in chunks of information much less than original sequence but more specialized for comparison. Three methods were developed using three different imaging functions. The first is based on identity of local sites of up to twelve nucleotides, the second is based on statistical homology of local 42 nucleotide fragments, and the third is based on the homology of 100-150 nucleotide fragments and models the comparison of restriction maps. Each of them requires the library of sequence images. The total size of such a library is less than the size of sequences stored in compressed form. The sequences are aligned allowing local homology searches. The method reduces total time for a similarity search about 100-fold. The programs can be easily included in any software, which allows user to define his own set of sequences. One of the programs is implemented within DNA-SUN software and is used in Institute of Molecular Genetics and Institute of Molecular Biology.

Amino Acid Sequence↗

The comparative amino acid sequences, substrate specificities and gene or cDNA nucleotide sequences of some prokaryote and eukaryote amidinotransferases: implications for evolution.

The amino acid sequences of the amidinotransferases and the nucleotide sequences of their genes or cDNA from four Streptomyces species (seven genes) and from the kidneys of rat, pig, human and human pancreas were compared. The overall amino acid and nucleotide sequences of the prokaryotes and eukaryotes were very similar and further, three regions were identified that were highly identical. Evidence is presented that there is virtually zero chance that the overall and high identity regions of the amino acid sequence similarities and the overall nucleotide sequence similarities between Streptomyces and mammals represent random match. Both rat and lamprey amidinotransferases were able to use inosamine phosphate, the amidine group acceptor of Streptomyces. We have concluded that the structure and function of the amidinotransferases and their genes has been highly conserved through evolution from prokaryotes to eukaryotes. The evolution has occurred with: (1) a high degree of retention of nucleotide and amino acid sequences; (2) a high degree of retention of the primitive Streptomyces guanine + cytosine (G + C) third codon position composition in certain high identity regions of the eukaryote cDNA; (3) a decrease in the specificities for the amidine group acceptors; and (4) most of the mutations silent in the regions suggested to code for active sites in the enzymes.

Amidinotransferases↗

Massive sequence comparisons as a help in annotating genomic sequences.

An all-by-all comparison of all the publicly available protein sequences from plants has been performed, followed by a clusterization process. Within each of the 1064 resulting clusters-containing sequences that are orthologous as well as paralogous-the sequences have been submitted to a pyramidal classification and their domains delineated by an automated procedure à la. This process provides a means for easily checking for any apparent inconsistency in a cluster, for example, whether one sequence is shorter or longer than the others, one domain is missing, etc. In such cases, the alignment of the DNA sequence of the gene with that of a close homologous protein often reveals (in 10% of the clusters) probable sequencing errors (leading to frameshifts) or probable wrong intron/exon predictions. The composition of the clusters, their pyramidal classifications, and domain decomposition, as well as our comments when appropriate, are available from http://chlora.infobiogen.fr:1234/PHYTOPROT.

Amino Acid Sequence↗

Intervening sequences in the ribosomal RNA genes of Ascaris lumbricoides: DNA sequences at junctions and genomic organization.

An rDNA size class in the genome of the nematode Ascaris lumbricoides is described which is interrupted by a 4.5-kb long intervening sequence located in the 26S coding region. This molecular form occurs in approximately 15 copies per haploid genome and amounts to approximately 5% of the total nuclear rDNA. Intervening sequences are present only in the 8.8-kb rDNA, but not in the 8.4-kb rDNA repeating units of A. lumbricoides. Cloning of the interrupted rDNA units revealed, in addition to the main 4.5-kb insertion, shorter intervening sequences of 4-kb and 119-bp length. Both shorter rDNA forms are present in the single copy range of the haploid genome. Sequence analyses of the intervening sequence/rDNA junctions show an identical right-hand junction for all of the three different rDNA forms. The two shorter intervening sequences are a coterminal subset of the right-hand end of the main 4.5-kb insertion, whereas all three insertions have a different left-hand junction with the coding region of rDNA. Each intervening sequence is flanked by a short direct repeat of variable length, being only once present in the uninterrupted rDNA. The intervening sequences of A. lumbricoides show striking similarity to the organization of type I insertion family in dipteran flies, even though they are inserted at different positions in the 26S coding region. Additional rDNA intervening sequences may be present outside of the rDNA cluster, but in not more than 15-20 homologous copies per haploid genome.

Animals↗

New approaches for innovations in sensitive Edman sequence analysis by design of a wafer-based chip sequencer.

In the last few years the development of new mass spectrometric techniques enabled fast and sensitive protein analysis by the introduction of mass spectrometry (MS) fingerprinting and MS/MS sequencing. For these methods mixtures of peptide fragments of the proteins can be employed, whereas the Edman degradation method requests purified peptides. On the other hand, Edman sequencing has the advantages that interpretation of the data is more simple, extended sequences can be derived, and reliable sequence information on unknown proteins is possible. Hence, Edman chemistry as an alternative technique to MS is still valuable. But higher sensitivity of the sequencers is needed in order to meet modern demands, e.g. in proteomics research. We designed a wafer-based chip sequencer for protein and peptide sequencing in the femto- to attomole range. The main advantage of our new design is the complete integration of dead volume free valves together with reactor and converter and volume-measuring loops within one wafer-based system. In this system aggressive chemicals and solvents for the Edman degradation can be delivered in sub-microliter amounts, which allows a considerable shortening of the degradation cycles. Further, we developed sensitive maintenance and tightness tests to prove a precise and reproducible delivery of the chemicals and the reduction of drying times as compared with conventional sequencers. Real parallel processing of several samples can easily be implemented. The system is designed to serve future needs in protein research.

Indicators and Reagents↗

The DNA sequences of cloned complex satellite DNAs from Hawaiian Drosophila and their bearing on satellite DNA sequence conservation.

A class of restriction endonuclease fragments near 185 bp in length and comprising approximately 20% of the genomes of 3 species of Hawaiian Drosophila has been cloned using bacteriophage M13. The nucleotide sequences of 14 clones have been determined and the variation between clones has been found to be due to deletions and base changes. Analyses of uncloned material show that the cloning system itself does not introduce the variation. The variation of the basic repeat within and between species is high; 15% due to deletions and 10% due to base changes. The Drosophila data are similar in many respects to both the 23 bp calf satellite results (Pech et al., 1979 b) and those from sequence analyses of the 170 bp primate restriction fragments (Rubin et al., 1979; Donehower et al., 1980, Wu and Manuelidis, 1980). The intraspecies level of base changes and deletions in the calf satellite approaches 25% as does that in the human/African green monkey/baboon comparisons. The between species variation in the primate group is near 35%. Direct sequencing methods thus reveal a widespread sequence heterogeneity in both invertebrate and mammalian satellite systems of long or short repeat length. This heterogeneity does not support the strict sequence conservation implied by the "library" hypothesis, which claims a functional role in speciation for the rigid conservation of satellite DNA sequences (Fry and Salser, 1977). Furthermore the Drosophila and primate data reveal that satellite DNAs can change rapidly, though nonrandomly, at the nucleotide sequence level in a relatively closely knit group such as the Hawaiian species, as well as in more distantly related species from amongst the primates. We draw two major conclusions. There is no universal attribute of satellite DNA sequence per se, the only biological variable to date being the amount of satellite DNA and its effect in the germ line. Many aspects of satellite DNA evolution conform to Kimura's (1979) concepts of neutrality.

Animals↗

Characterisation of a human Y chromosome repeated sequence and related sequences in higher primates.

The human Y chromosome carries 2000 copies of a tandemly repeated sequence, 2.47 kb long, which constitutes about 20% of the DNA of this chromosome. These sequences are localised on the tip of the long arm of the Y chromosome. Related sequences are present in DNA of females with a related but distinguishable restriction pattern. These autosomal sequences are distributed in tandem arrays on a number of autosomes. Related sequences are also present in gorilla and chimpanzee. In gorilla they resemble the human sequences in their restriction map but are not found on the Y chromosome whereas in chimpanzee the related sequences behave as a 'dispersed' repeat. Changes in the level of methylation of this sequence in different tissues of human males can be detected with the lowest levels found in sperm and placental DNA.

Animals↗

Effectiveness of measures requiring and not requiring prior sequence alignment for estimating the dissimilarity of natural sequences.

Various measures of sequence dissimilarity have been evaluated by how well the additive least squares estimation of edges (branch lengths) of an unrooted evolutionary tree fit the observed pairwise dissimilarity measures and by how consistent the trees are for different data sets derived from the same set of sequences. This evaluation provided sensitive discrimination among dissimilarity measures and among possible trees. Dissimilarity measures not requiring prior sequence alignment did about as well as did the traditional mismatch counts requiring prior sequence alignment. Application of Jukes-Cantor correction to singlet mismatch counts worsened the results. Measures not requiring alignment had the advantage of being applicable to sequences too different to be critically alignable. Two different measures of pairwise dissimilarity not requiring alignment have been used: (1) multiplet distribution distance (MDD), the square of the Euclidean distance between vectors of the fractions of base signlets (or doublets, or triplets, or ...) in the respective sequences, and (2) complements of long words (CLW), the count of bases not occurring in significantly long common words. MDD was applicable to sequences more different than was CLW (noncoding), but the latter often gave better results where both measures were available (coding). MDD results were improved by using longer mutliplets and, if the sequences were coding, by using the larger amino acid and codon alphabets rather than the nucleotide alphabet. The additive least squares method could be used to provide a reasonable consensus of different trees for the same set of species (or related genes).

Animals↗

The complete sequence of the rice (Oryza sativa L.) mitochondrial genome: frequent DNA sequence acquisition and loss during the evolution of flowering plants.

The entire mitochondrial genome of rice (Oryza sativa L.), a monocot plant, has been sequenced. It was found to comprise 490,520 bp, with an average G+C content of 43.8%. Three rRNA genes, 17 tRNA genes and five pseudo tRNA sequences were identified. In addition, eleven ribosomal protein genes and two pseudo ribosomal protein genes were found, which are homologous to 13 of the 16 genes for ribosomal proteins in the mitochondrial genome of the liverwort (Marchantia polymorpha). A greater degree of variation in terms of presence/absence and integrity of genes was observed among the ribosomal protein genes and tRNA genes of rice, Arabidopsis and sugar beet. Transcription and post-transcriptional modification (RNA editing) in the rice mitochondrial sequence were also examined. In all, 491 Cs in the genomic DNA were converted to Ts in cDNA. The frequency of RNA editing differed markedly depending upon the ORF considered. Sequences derived from plastid and nuclear genomes make up 6.3% and 13.4% of the mitochondrial genome, respectively. The degree of conservation of plastid sequences in the mitochondrial genome ranged from 61% to 100%, suggesting that sequence migration has occurred very frequently. Three plastid DNA fragments that were incorporated into the mitochondrial genome were subsequently transferred to the nuclear genome. Nineteen fragments that were similar to transposon or retrotransposon sequences, but different from those found in the mitochondrial genomes of dicots, were identified. The results indicate frequent and independent DNA sequence flow to and from the mitochondrial genome during the evolution of flowering plants, and this may account for the range of genetic variation observed between the mitochondrial genomes of higher plants.

Biological Evolution↗

Sequence analysis of adenovirus DNA: complete nucleotide sequence of the spliced 5' noncoding region of adenovirus 2 hexon messenger RNA.

The complete nucleotide sequence of the 5' noncoding region of the adenovirus 2 hexon messenger RNA has been established by sequence analysis of reverse transcripts. Such transcripts were generated by extension of specific single-stranded DNA primers with reverse transcriptase after hybridization to purified hexon mRNA. The total length of the 5' noncoding region was determined to be 240 nucleotides, of which the spliced tripartite leader sequence contributes 202 nucleotides including the terminal m7G. The sizes of the different segments of the tripartite leader were estimated by comparing the established mRNA sequence with the genomic sequences for the first and third leader segments, and were found to be 42 nucleotides for the first segment, 71 nucleotides for the second and 89 nucleotides for the third. The estimates are ambiguous, however, due to the presence of tandemly repeated sequences at both ends of the intervening sequence between the third leader segment and the body of the hexon mRNA. The sequence of the leader allows the formation of hydrogen-bonded interactions with the 3' end of 18S ribosomal RNA near the capped 5' end and also close to the initiator AUG.

Adenoviruses, Human↗

Poly(dG)-poly(dC) sequences, under torsional stress, induce an altered DNA conformation upon neighboring DNA sequences.

Supercoiled plasmid DNAs (at bacterial superhelical density) harboring the homopurine-homopyrimidine sequence, poly(dG)-poly(dC), were reacted with bromoacetaldehyde (BAA), a reagent that reacts with unpaired DNA bases. Not only did the poly(dG)-poly(dC) sequence react with BAA but, surprisingly, neighboring sequences located 3' to the contiguous G sequences also reacted. The altered conformation in the poly(dG)-poly(dC) sequence and in the neighboring sequence occurred in the same supercoiled plasmid DNA molecule. Furthermore, the occurrence of an "unpaired" conformation in the neighboring sequence is strictly due to a positional effect, since it is observed when the poly(dG)-poly(dC) segment is adjacent to a variety of neighboring sequences.

Acetaldehyde↗

Interlaboratory concordance of DNA sequence analysis to detect reverse transcriptase mutations in HIV-1 proviral DNA. ACTG Sequencing Working Group. AIDS Clinical Trials Group.

Thirteen laboratories evaluated the reproducibility of sequencing methods to detect drug resistance mutations in HIV-1 reverse transcriptase (RT). Blinded, cultured peripheral blood mononuclear cell pellets were distributed to each laboratory. Each laboratory used its preferred method for sequencing proviral DNA. Differences in protocols included: DNA purification; number of PCR amplifications; PCR product purification; sequence/location of PCR/sequencing primers; sequencing template; sequencing reaction label; sequencing polymerase; and use of manual versus automated methods to resolve sequencing reaction products. Five unknowns were evaluated. Thirteen laboratories submitted 39043 nucleotide assignments spanning codons 10-256 of HIV-1 RT. A consensus nucleotide assignment (defined as agreement among > or = 75% of laboratories) could be made in over 99% of nucleotide positions, and was more frequent in the three laboratory isolates. The overall rate of discrepant nucleotide assignments was 0.29%. A consensus nucleotide assignment could not be made at RT codon 41 in the clinical isolate tested. Clonal analysis revealed that this was due to the presence of a mixture of wild-type and mutant genotypes. These observations suggest that sequencing methodologies currently in use in ACTG laboratories to sequence HIV-1 RT yield highly concordant results for laboratory strains; however, more discrepancies among laboratories may occur when clinical isolates are tested.

Codon↗

Two-step high resolution sequence-based HLA-DRB typing of exon 2 DNA with taxonomy-based sequence analysis allele assignment.

A two-step high resolution sequence-based DRB typing method was developed. The system needs only one polymerase chain reaction (PCR) to type all functional DRB alleles of a given individual. It uses a pair of generic PCR primers to amplify exon 2 DNA of all functional DRB genes and a first-step taxonomy-based sequence analysis (FSTBSA) method to assign allele groups after sequencing the PCR products with a generic primer. In the second step, group-specific primers are used to sequence the same PCR products and a taxonomy-based sequence analysis (TBSA) is used to assign alleles. Thus, both low and high resolution DRB typing can be done with PCR amplified exon 2 DNA from a single PCR reaction. Correct allele group assignment by FSTBSA was confirmed by sequencing the PCR products with group-specific primers and correctly assigned all 158 DNA samples including 34 samples pre-typed by PCR-sequence-specific primer or PCR-sequence-specific oligonucleotide probe. FSTBSA correctly assigned 116 heterozygous combinations of 81 DRB1-DRB3/4/5 haplotypes. Sixty-seven DRB1, 6 DRB3, 1 DRB4, and 3 DRB5 alleles were identified in this study. TBSA successfully resolved all heterozygous allele combinations including 31 heterozygous combinations of 33 alleles of DRB1*03, 08, 11, 12, 13, and 14 allele groups, and six heterozygous combinations of six DRB3 alleles.

Alleles↗