Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Genome-wide detection of alternative splicing in expressed sequences using partial order multiple sequence alignment graphs.

We present a method for high-throughput alternative splicing detection in expressed sequence data. This method effectively copes with many of the problems inherent in making inferences about splicing and alternative splicing on the basis of EST sequences, which in addition to being fragmentary and full of sequencing errors, may also be chimeric, misoriented, or contaminated with genomic sequence. Our method, which relies both on the Partial Order Alignment (POA) program for constructing multiple sequence alignments, and its Heaviest Bundling function for generating consensus sequences, accounts for the real complexity of expressed sequence data by building and analyzing a single multiple sequence alignment containing all of the expressed sequences in a particular cluster aligned to genomic sequence. We illustrate application of this method to human UniGene Cluster Hs.1162, which contains expressed sequences from the human HLA-DMB gene. We have used this method to generate databases, published elsewhere, of splices and alternative splicing relationships for the human, mouse and rat genomes. We present statistics from these calculations, as well as the CPU time for running our method on expressed sequence clusters of varying size, to verify that it truly scales to complete genomes.

Alternative Splicing↗

Cloned mRNA sequences for two types of embryonic myosin heavy chains from chick skeletal muscle. I. DNA and derived amino acid sequence of light meromyosin.

Two myosin heavy chain cDNA clones (251 and 110), constructed from chick embryonic skeletal muscle mRNA, were subjected to extensive DNA sequence analysis. A complete description of the DNA sequence of clone 251 was obtained. This 1.5-kilobase pair cDNA sequence specified the COOH-terminal 439 amino acids of the myosin heavy chain, and included the entire 3' nontranslated region. The translated and 3' nontranslated sequences were purine- (64%) and AT-(71%) rich, respectively. The derived amino acid sequence of clone 251 correlated well with sequences obtained by direct amino acid sequencing of adult rabbit back muscle myosin heavy chain protein (87% homology), as well as with cloned myosin heavy chain sequences from other species. Comparison of clone 251 with a partial DNA sequence of clone 110 revealed significant structural differences both in the translated, and 3' nontranslated regions. This data indicates that these two clones represent two distinct myosin heavy chain genes. The protein sequence specified by clone 251 corresponds to the light meromyosin portion of the myosin heavy chain rod. These sequences, like other myosin heavy chain rod sequences, are alpha-helical and exhibit 7- and 28-residue periodicities in the linear distribution of nonpolar, and basic and acidic amino acids, respectively.

Amino Acid Sequence↗

Rapid sequencing of the p53 gene with a new automated DNA sequencer.

p53 is the most commonly mutated gene in human cancers. Approximately 90% of the p53 gene mutations are localized between domains encoding exons 5 to 8. Sequencing methods currently available are tedious and time-consuming and are not suitable for routine laboratory testing. In an effort to identify a simple and rapid sequencing method, we analyzed 16 preselected breast tumors and 18 preselected ovarian tumors, using a newly developed automated DNA sequencer. p53 gene mutations had been previously identified in these tumors, using a conventional automated sequencing procedure. Exons 5 to 8 were amplified by PCR, and the PCR products were subsequently subjected to cycle sequencing with the Sanger chain termination method, using Cy5.5-labeled primers. The sequencing mixture was then resolved on a newly developed automated DNA sequencer that can sequence approximately 300 bases of DNA in 30 min. Of these 16 breast tumors, two had mutations in exon 5, four in exon 6, three in exon 7, and three in exon 8. Of the 18 ovarian tumors, two had mutations in exon 5, five in exon 6, two in exon 7, and three in exon 8. In all cases, we identified the same mutations by both the new and the conventional sequencing procedures. Most mutations affected an arginine codon. These data demonstrate that the new method has the capability to provide accurate sequencing information in a fraction of the time and labor in comparison with current automated sequencing techniques. When such procedures are used, DNA sequencing may become a routine tool for identifying clinically important mutations for diagnosis and prognosis of patients with genetic, malignant, infectious, and other diseases.

Autoanalysis↗

Comparison of DNA sequences with protein sequences.

The FASTA package of sequence comparison programs has been expanded to include FASTX and FASTY, which compare a DNA sequence to a protein sequence database, translating the DNA sequence in three frames and aligning the translated DNA sequence to each sequence in the protein database, allowing gaps and frameshifts. Also new are TFASTX and TFASTY, which compare a protein sequence to a DNA sequence database, translating each sequence in the DNA database in six frames and scoring alignments with gaps and frameshifts. FASTX and TFASTX allow only frameshifts between codons, while FASTY and TFASTY allow substitutions or frameshifts within a codon. We examined the performance of FASTX and FASTY using different gap-opening, gap-extension, frameshift, and nucleotide substitution penalties. In general, FASTX and FASTY perform equivalently when query sequences contain 0-10% errors. We also evaluated the statistical estimates reported by FASTX and FASTY. These estimates are quite accurate, except when an out-of-frame translation produces a low-complexity protein sequence. We used FASTX to scan the Mycoplasma genitalium, Haemophilus influenzae, and Methanococcus jannaschii genomes for unidentified or misidentified protein-coding genes. We found at least 9 new protein-coding genes in the three genomes and at least 35 genes with potentially incorrect boundaries.

Amino Acid Sequence↗

Determination of the position of the boundaries of the terminal repetitive sequences within the genome of molluscum contagiosum virus type 1 by DNA nucleotide sequence analysis.

The repetitive DNA sequences of the genome of Molluscum contagiosum virus type 1 (MCV-1) have been localized within the terminal regions of the viral genome corresponding to the BamHI MCV-1 DNA fragments B (18 kbp; 0 to 0.095 map units (m.u.)) and E (10.5 kbp; 0.944 to 1 m.u.). The fine mapping of these particular regions of the genome of MCV-1 revealed that the boundaries of the terminal repetitive DNA sequences of the viral genome are located within the DNA sequences of the HindIII MCV-1 DNA fragments K (3.8 kbp; 0.014 to 0.036 m.u.) and J1 (4.1 kbp; 0.962 to 0.985 m.u.). The exact position of the boundary of the repetitive DNA sequences was determined by DNA nucleotide sequencing. The HindIII DNA fragments K and J1 compose 3859 and 4107 bp, respectively. The DNA sequences of HindIII MCV-1 DNA fragment K possess repetitive DNA sequences between the nucleotide positions 1 and 1675 which are homologous to the inverted and complementary DNA sequences of the HindIII MCV-1 DNA fragment J1 between the nucleotide positions 2437 and 4107 (1670 bp). The degree of DNA sequence homology detected between the repetitive DNA sequences in the HindIII DNA fragments K and J1 of the viral genome was found to be 98%. The number of open reading frames (ORFs) detected by the analysis of the DNA sequences of the HindIII MCV-1 DNA fragments K and J1 was found to be 14 (70 to 219 amino acid residues) and 11 (70 to 365 amino acid residues), respectively.

Base Sequence↗

Molecular cloning, DNA sequence analysis, and expression of cDNA sequence of RNA genomic segment 6 (S6) that encodes a viral outer capsid protein of threadfin aquareovirus (TFV).

The genome segment 6 (S6) of threadfin reovirus (TFV) was cloned and sequenced. The entire S6 nucleotide sequence is 2056 bp long with an open reading frame that encodes a protein of 653 amino acids. Sequence analysis of the TFV S6 genome revealed that the 5'-terminal sequence, GTTTTA and the 3'-terminal sequence, ATTCATC of the plus strand is common to other genome segments of TFV. The pentanucleotide, TCATC, at the 3'-terminal of the plus strand was also conserved in other reported isolates of Aquareovirus such as chum salmon reovirus (CSV), striped bass reovirus (SBR), grass carp reovirus (GCRV) and golden shiner reovirus (GSV) as well as to the 10 genome segments of mammalian reovirus (MRV). Blast results indicated that the TFV S6 gene segment sequence had high identity towards the CSV S6 gene sequence, which codes for the CSV outer coat protein. This implied that the TFV S6 gene segment codes for an outer capsid protein (OCP) of the virus. Amino acid sequence analysis of this TFV OCP sequence revealed the presence of a putative conserved asparagine-proline (Asn-Pro) protease cleavage site, which was found in all reported isolates of Aquareovirus as well as in the MRV mu1 protein. N-terminal sequencing of the corresponding S6 native protein obtained from purified TFV particles verified the presence of this cleavage site. Phylogenetic analysis of the TFV S6 protein revealed that TFV was closely related to CSV, from Aquareovirus species, ARV-A. Cloning of the TFV S6 gene sequence into an Escherichia coli expression host produced a recombinant protein that corresponded to the predicated size of the OCP of TFV. Immunization of mice using this recombinant outer capsid protein (rOCP) revealed that the protein was able to elicit an antibody response, thus indicating that the rOCP of TFV was immunogenic.

Amino Acid Sequence↗

Nucleotide sequence of a mouse lamin A cDNA and its deduced amino acid sequence.

We have cloned and determined the nucleotide sequence of a mouse lamin A cDNA. The clone contained the C-terminal two-thirds of the lamin A coding sequence and a 3' untranslated sequence with a poly(A) stretch. As has been reported for human lamin A/C cDNAs, a large part of the 5' sequence of our mouse lamin A clone was essentially identical with a previously reported mouse lamin C cDNA sequence, and the deduced C-terminal amino acid sequence shared strong homology with the human lamin A sequence. A putative deduced amino acid sequence for mouse lamin A, which was derived from our sequence and the published lamin C sequence, was 665 amino acids long. The degree of overall homology to the human sequence was more than 95%, and relatively more variation was scattered in the C-terminal lamin A-specific region.

Amino Acid Sequence↗

Molecular cloning and sequencing of the MTV-1 LTR: evidence for a LTR sequence alteration.

The vertically transmitted Mtv-1 provirus is the primary causative factor of mammary neoplasia in certain C3Hf strains that lack the horizontally transmitted mouse mammary tumor virus (MMTV). The studies here report the molecular cloning of the germ line 4.5 kb Mtv-1 3' EcoRI fragment and sequencing of the 3' Mtv-1 LTR. The Mtv-1 LTR sequence is closely related to the 5' Mtv-11 LTR sequence also reported here, as well as to known Mtv-8 and MMTV LTR sequences in the portion of MMTV and Mtv-8 LTRs previously demonstrated to contain transcriptional regulatory sequences. A 91 bp unique sequence region, Mtv-1 bp 862 to 952, exists in the Mtv-1 LTR, which is upstream of the sequence homology with the MMTV transcriptional regulatory domain. The Mtv-1 unique sequence region is distinct from a 117 bp sequence, bp 862 to 978, in the Mtv-11 LTR sequence as well as reported Mtv-8 and MMTV LTR sequences, and is present in the germ line Mtv-1 5' and 3' LTR-containing restriction fragments. S1 nuclease mapping experiments of C3Hf/Se mammary tumor poly(A) RNA with the cloned Mtv-1 and Mtv-11 LTRs exhibited a specific set of S1 protected fragments demonstrating that Mtv transcripts which accumulate in C3Hf spontaneous mammary tumors are encoded by the Mtv-1 provirus.

Animals↗

Inverted repeat regions of Marek's disease virus DNA possess a structure similar to that of the a sequence of herpes simplex virus DNA and contain host cell telomere sequences.

The genomic structure of Marek's disease virus (MDV) is similar to those of the alphaherpesviruses herpes simplex virus (HSV) types 1 and 2. Sequence analysis of the junction region between the long component (L) and the short component (S) revealed the existence of an a-like sequence, similar in structure to the a sequence of HSV-1. Further study revealed that the MDV genome contains five copies of the a-like sequence within the long terminal repeat region as well as in the short terminal repeat region. The junction between the L and S components was found to contain 10 copies of the a-like sequence. Within the a-like sequence, a structure homologous to the DR2 of HSV was found to contain 17 copies of the telomeric sequence, GGGGTTA. There appears to be little to no sequence homology between the HSV a sequence and the MDV a-like sequence; however, the strong physical homology to its counterpart in HSV-1 suggests that the MDV a-like sequence may have the same functional homology (the domain for cleavage/packaging of the DNA into the viral capsids and for genomic inversion) as well.

Animals↗

Primary structure of the H-2Db alloantigen. II. Additional amino acid sequence information, localization of a third site of glycosylation and evidence for K and D region specific sequences.

The complete amino acid sequence of the CNBr fragment comprising residues 229-284 of the murine major histocompatibility complex antigen H-2Db has been determined using radiochemical methodology. The sequence was determined by N-terminal sequence analysis of the intact CNBr fragment and by sequence determinations of peptides derived from this fragment by trypsin and staphylococcal V8 protease cleavage. In addition to the amino acid assignments for H-2Db, it was possible to assign the linkage position of the third N-linked glycosyl unit to the asparagine at residue 256. Additional amino acid sequence assignments have also been made for three other CNBr fragments that span residues 99-138, 139-228, and 308-331 of the H-2Db molecule. The total protein sequence information available (222 of 338 residues) agrees in every comparable position with the protein sequence derived from the cDNA clone (pH203) isolated by Reyes and co-workers (1982b), which strongly suggests that this clone encodes H-2Db. Combination of the protein sequence with that deduced from the cDNA clone provides the complete H-2Db protein sequence. Comparison of this sequence with other available protein sequence information for murine class I molecules has revealed protein sequences that may be unique to either K or D region molecules.

Amino Acid Sequence↗

Population data for 101 Austrian Caucasian mitochondrial DNA d-loop sequences: application of mtDNA sequence analysis to a forensic case.

The sequence of the two hypervariable segments of the mitochondrial DNA (mtDNA) control region was generated for 101 random Austrian Caucasians. A total of 86 different mtDNA sequences was observed, where 11 sequences were shared by more than 1 individual, 7 sequences were shared by 2 individuals and 4 sequences were shared by 3 individuals. One of the four most common mtDNA sequences in Austrians is also the most common sequence in both U.S. and British Caucasians, found in approximately 3.0% of Austrians, 4.0% of British, and 3.9% of U.S. Caucasians. Of the remaining three common Austrian sequences, one was not observed in either U.S. or British Caucasians. However, three British Caucasians exhibited a similar sequence type. Therefore, this particular cluster of sequence polymorphisms may represent a common "European" mtDNA sequence type. In general, Austrian Caucasians show little deviation from other Caucasian databases of European descent. Finally, mtDNA sequence analysis was applied to a forensic case, where hairs found at a crime scene matched the control hairs from the suspect.

Austria↗

A 9.6 kb intervening sequence in D. virilis rDNA, and sequence homology in rDNA interruptions of diverse species of Drosophila and other diptera.

A large proportion of the 28S ribosomal RNA genes in Drosophila virilis are interrupted by a DNA sequence 9.6 kilobase pairs long. As regards both its presence and its position in the 28S gene (about two thirds of the way in), the D. virilis rDNA intervening sequence is similar to that found in D. melanogaster rDNA, but lengths differ markedly between the two species. Degrees of nucleotide sequence homology have been detected bewteen rDNA interruptions of the two species. This homology extends to putative rDNA intervening sequences in diverse higher diptera (other Drosophila species, the house fly and the flesh fly), but hybridization of cloned D. melanogaster and D. virilis rDNA interruption segments to DNA of several lower diptera has been negative. As is the case with melanogaster rDNA interruptions, segments of the virilis rDNA intervening sequence hybridize with non-rDNA components of the virilis genome, and interspecific homology may involve these non-rDNA sequences as well as rDNA interruptions. There is, however, evidence from buoyant density fractionation of DNA that the distributions of interruption-related sequences are distinct in D. melanogaster and D. virilis genomes. Moreover, thermal denaturation studies have indicated differing extents of homology between hybridizable sequences in D. virilis DNA and different segments of the D. melanogaster rDNA intervening sequence. We infer from our studies that rDNA intervening sequences are prevalent among higher diptera; that in the course of the evolution of these organisms, elements of the intervening sequences have been moderately to highly conserved; and that this conservation extends in at least two distantly related species of Drosophila to similar sequences found elsewhere in the genomes.

Animals↗

Splice point sequence and transcripts of the intervening sequence in the mitochondrial 21S ribosomal RNA gene of yeast.

By S1 nuclease mapping we have located the intervening sequence in the large ribosomal RNA gene of Saccharomyces cerevisiae omega+ strains 570 bp from the 3' end of the rRNA gene. No intervening sequence was detected at this position in S. carlsbergensis, but the sequences of the mature 21S rRNAs of these two strains appear to be identical in this region. By comparing the DNA sequence of the region of the intervening sequence in an omega+ strain with the corresponding sequence in S. carlsbergensis, we have determined the splice points of the 21S rRNA gene. These sequences show no homology with splice points in nuclear and viral genes or with the splice points in the chloroplast 23S rRNA gene of Chlamydomonas. The external borders of the splice points have a complementary sequence in the intervening sequence. The largest transcript hybridizing with the probe of the intervening sequence has a size corresponding to that expected for an rRNA precursor still containing the intervening sequence; the smallest transcript corresponds in size to the intervening sequence itself.

Base Sequence↗

Simultaneous sequencing of multiple polymerase chain reaction products and combined polymerase chain reaction with cycle sequencing in single reactions.

DNA sequencing is considered the gold standard for nucleic acid identification and mutation detection. However, sequencing is labor intensive because it requires previous amplification and only a single sequence is analyzed at a time. We developed two novel strategies that substantially improve DNA sequencing. The first allows multiple polymerase chain reaction (PCR) products to be sequenced in a single sequencing reaction and analyzed simultaneously in a single lane or capillary. Simultaneous sequencing by this method, designated "SimulSeq," can provide either simultaneous single-direction sequencing of multiple genes or simultaneous forward and reverse sequencing from a single gene. In the second approach, designated "AmpliSeq," we demonstrate a technique combining PCR amplification and sequencing in a single reaction that is analyzed in a single lane or capillary. We demonstrate combined PCR with short bidirectional sequencing, and combined PCR with unidirectional sequencing. We anticipate that these methods will have utility in research and clinical settings where panels of mutations or large numbers of samples are be analyzed and/or when turnaround time is critical.

Base Sequence↗