Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “DNA sequence analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

Complexity: an internet resource for analysis of DNA sequence complexity.

The search for DNA regions with low complexity is one of the pivotal tasks of modern structural analysis of complete genomes. The low complexity may be preconditioned by strong inequality in nucleotide content (biased composition), by tandem or dispersed repeats or by palindrome-hairpin structures, as well as by a combination of all these factors. Several numerical measures of textual complexity, including combinatorial and linguistic ones, together with complexity estimation using a modified Lempel-Ziv algorithm, have been implemented in a software tool called 'Complexity' (http://wwwmgs.bionet.nsc.ru/mgs/programs/low_complexity/). The software enables a user to search for low-complexity regions in long sequences, e.g. complete bacterial genomes or eukaryotic chromosomes. In addition, it estimates the complexity of groups of aligned sequences.

Algorithms↗

Complete mitochondrial DNA sequence and amino acid analysis of the cytochrome C oxidase subunit I (COI) from Aedes aegypti.

The complete sequence of the yellow fever mosquito, Aedes aegypti, mitochondrial cytochrome c oxidase subunit 1 gene has been identified. The nucleotide sequence codes for a 512 amino acid peptide. The AeCOI sequence is A + T rich (68.6%) and the codon usage is highly biased toward a preference for A- or T-ending triplets. The A. aegypti COI peptide shows high homology, up to 93% identity, with several other insect sequences and a phylogenetic analysis indicates that the A. aegypti sequence is closely related to two other mosquito species, Anopheles gambiae and A. quadrimaculatus. Comparisons of the nucleotide sequence for four A. aegypti laboratory strains revealed single nucleotide polymorphisms, with 25 nucleotide sites showing SNPs between strains. All SNPs occurred as synonymous transitions such that the peptide sequence is conserved among A. aegypti strains. RT-PCR analysis showed that COI is expressed at similar levels in all developmental stages and tissues.

AT Rich Sequence↗

Molecular cloning and characterization of the bovine tuftelin gene.

The bovine tuftelin gene was cloned and its structure determined by DNA sequence analysis and comparison to that of bovine tuftelin cDNA. The analyses demonstrated that the cDNA contains a 1014-bp open reading frame encoding a protein of 338 residues with a calculated mol. wt of 38,630 and an isoelectric point of 5.85. These results differ from those previously published, (1991) which contained a different conceptual amino acid sequence for the carboxy terminal region and identified a different termination codon. The protein does not appear to share homology or domain motifs with any other known protein. The gene consists of 13 exons ranging in size from 66 to 1531 bp, the latter containing the encoded carboxyterminal and 3' untranslated regions. The exons are embedded in more than 28 kbp of genomic DNA. Codons are generally not divided at exon/intron borders. Several alternatively spliced transcripts were identified by DNA sequence analysis of the isolated products produced by reverse transcriptase/polymerase chain reaction.

Alternative Splicing↗

[PCR-single-strand conformation (SSCP), DNA direct sequencing analysis in detecting mutation in exon 2 of g6pd gene].

Hereditary glucose-6-phosphate dehydrogenase (G6PD) in red blood cell was one of the most common genetic diseases in South China. The research data of G6PD gene showed that at least 6 point mutations were responsible for various G6PD variants in Chinese. For developing a rapid, sensitive, and effective method to detect point mutation, we applied single-strand conformation (SSCP) analysis for detection of mutation in exon 2 of G6PD gene of 20 cases of G6PD variants. Four of them were found that mobility shift band in one of two single strands DNA is slower than other individuals. PCR direct sequencing for these 4 samples were given a base substitution (T to D) at nucleotide 95 of cDNA. The results indicated that this technique is very simple, sensitive, and useful over other methods of detecting point mutation. The experimental conditions of PCR-SSCP and features of this mutation were discussed at the same time.

Adult↗

A pathogenicity gene isolated from the pPATH plasmid of Erwinia herbicola pv. gypsophilae determines host specificity.

The host range of the gall-forming bacterium Erwinia herbicola pv. gypsophilae (Ehg) is restricted to the gypsophila plant whereas E. herbicola pv. betae (Ehb) incites galls on beet as well as gypsophila. The pathogenicity of Ehg and Ehb was previously shown to be dependent on a plasmid (pPATH). Transposon mutagenesis was used to generate mutants on the cosmid pLA150 of the pPATH from Ehg824-1. A cluster of nonpathogenic mutations flanked by two IS1327 elements was identified on a 3.2-kb NdeI DNA fragment. All mutants were restored to pathogenicity by complementation in trans with the wild-type Ehg DNA. DNA sequence analysis of the 3.2-kb NdeI fragment revealed a single open reading frame (ORF) of 2 kb as well as a potential ribosome binding site and a putative hrp box upstream to the ORF. The ORF had no significant homology to known genes. Southern analysis also revealed the presence of DNA sequences that hybridized to the ORF in the beet pathovar Ehb4188. This gene was isolated and sequenced. Marker exchange mutants generated in the ORF of Ehb eliminated the pathogenicity of Ehb on gypsophila but fully retained its pathogenicity on beet. Since the putative gene appeared to encode a host-specific virulence factor for gypsophila it was designated as hsvG.

Amino Acid Sequence↗

Inactivation of p53 gene expression by an insertion of Moloney murine leukemia virus-like DNA sequences.

Analysis of Abelson murine leukemia virus-transformed L12 cells which lack the p53 cellular encoded tumor antigen revealed alterations in the p53-specific genomic DNA sequences. The active p53 gene, usually contained in a 16-kilobase EcoRI DNA fragment of p53 producer cells, went through major alterations leading to the appearance of a substantially larger 28.0-kilobase p53-specific EcoRI fragment. Detailed restriction enzyme analysis, with genomic probes spanning throughout the whole active p53 gene, indicated that the L12 p53 altered gene contains all the exons and principal introns of the normal p53 16.0-kilobase gene. However, its structure was interrupted by the integration of a novel DNA segment into the noncoding intervening sequences of the first p53 intron. Analysis of the inserted sequences revealed close homology to Moloney murine leukemia virus. This Moloney leukemia murine virus-like particle resides in a 5' to 3' transcriptional orientation, similar to the p53 gene, permitting the transcription of aberrant fused mRNA molecules detected in these cells.

Animals↗

A multiple-capillary electrophoresis system for small-scale DNA sequencing and analysis.

A five-capillary system has been developed for DNA sequencing and analysis. The post-column fluorescence detector is based on a sheath-flow cuvette. The instrument provides uniform and continuous illumination of the samples. The cuvette virtually eliminates cross-talk in the fluorescence signal between capillaries. Discrete single-photon counting avalanche photodiodes provide high efficiency light detection. The instrument has detection limits (3sigma) of 130 +/- 30 fluorescein molecules injected onto each capillary. Over 650 bases of sequence at 98.8% accuracy were generated in 100 min at 50 degrees C from M13mp18. Separation and detection of short tandem repeats proved efficient and accurate with the use of internal standards for direct comparison of migration times between capillaries.

Adult↗

A pathway-specific transcriptional regulatory gene for nikkomycin biosynthesis in Streptomyces ansochromogenes that also influences colony development.

DNA sequence analysis of a 7.5 kb XhoI DNA fragment from the region flanking the nikkomycin biosynthesis gene cluster in Streptomyces ansochromogenes revealed one 3.3 kb open reading frame (ORF), designated sanG. The deduced product of sanG (1061 amino acids), which is similar to PimR of Streptomyces natalensis, contains an OmpR-like DNA binding domain in its N-terminal portion and A- and B-type nucleotide binding motifs in the middle of the protein. Disruption of sanG abolished nikkomycin biosynthesis, reduced sporulation and led to brown pigment accumulation. All aspects of this complex phenotype were complemented by a single copy sanG which was integrated into the chromosome. The introduction of multiple copies of sanG resulted in increased nikkomycin production. S1 mapping results indicated that sanG is transcribed from at least three promoters (P1, P2 and P3), P1 being strongly upregulated when production of nikkomycins starts. Two putative transcription units for nikkomycin biosynthesis, starting from sanN and sanO, are dependent on the expression of sanG, whereas a putative transcription unit starting from sanF was not regulated by sanG. These results suggested that sanG encodes a transcriptional activator important for nikkomycin biosynthesis that, unusually, also has pleiotropic effects on secondary metabolism and development.

Amino Acid Sequence↗

Construction of the enhanced yellow fluorescent protein expression vector carrying IFN-gamma gene.

PURPOSE: To construct the enhanced yellow fluorescent protein (EYFP) vector carrying interferon-gamma gene (ifn-gamma) in order to provide an ideal reporter in the expression of ifn-gamma and location of protein in vitro and in vivo. METHOD: According to the nucleotide sequence of ifn-gamma gene, a pair of oligonucleotides was designed as primer whose two end contained nucleotide sequence of EcoR V and Not I restriction endonuclease respectively. The gene encoding for inf-gamma was amplified using PCR technqiue. After the PCR product was retrieved and purified, it was digested with EcoR V and Not I restriction endonuclease, and then cloned into the plasmid pIRES-EYFP. The recombinant plasmid pIRES-EYFPIFN-gamma was identified by restriction endonuclease enzyme analysis and DNA sequence analysis. RESULTS: The ifn-gamma was successfully amplified and verified by partial DNA sequence analysis. The recombinant plasmid was correctly screened. CONCLUSION: The EYFP expression vector carrying ifn-gamma gene was successfully established. This research work has formed a base for monitoring the ifn-gamma gene expression and protein position in living cells.

Bacterial Proteins↗

Cloning of the Retinoblastoma cDNA from the Japanese medaka (Oryzias latipes) and preliminary evidence of mutational alterations in chemically-induced retinoblastomas.

We have cloned a medaka homolog of the human retinoblastoma (Rb) susceptibility gene. The medaka Rb cDNA encodes a predicted protein of 909 amino acids. DNA sequence analysis with other vertebrate Rb sequences demonstrates that the medaka Rb cDNA is highly conserved in regions of functional importance. An antibody raised against an epitope of the human pRb recognizes the protein product of the medaka Rb gene, detecting a 105 kDa protein in all tissues examined and at differential levels for the stages of embryonic development studied. The sequence reported herein, combined with the high degree of conservation observed in critical domains, has also facilitated a preliminary investigation of the molecular etiology of chemically-induced retinoblastoma. The mutational alterations characterized suggest that medaka may provide a novel model and, thus, provide additional insight into the human retinoblastoma condition.

Amino Acid Sequence↗

Phylogenetic classification of serotype III group B streptococci on the basis of hylB gene analysis and DNA sequences specific to restriction digest pattern type III-3.

Previous work divided serotype III group B streptococci (GBS) into 3 major phylogenetic lineages (III-1, III-2, and III-3) on the basis of bacterial DNA restriction digest patterns (RDPs). Most neonatal invasive disease was caused by III-3 strains, which implies that III-3 strains are more virulent than III-2 or III-1 strains. In the current studies, all RDP III-3 and III-1 strains expressed hyaluronate lysase activity; however, all III-2 strains lack hyaluronate lysase activity, because the gene that encodes hyaluronate lysase, hylB, is inactivated by IS1548. Subtractive hybridization was used to identify 9 short DNA sequences that are present in all the III-3 strains but not in any of the III-2 or III-1 strains. With 1 exception, these III-3-specific sequences were not detected in nonserotype III GBS. These data further validate the RDP-based subclassification of GBS and suggest that lineage-specific genes will be identified, which account for the differences in virulence among the lineages.

Blotting, Southern↗

Exceptional motifs in different Markov chain models for a statistical analysis of DNA sequences.

Identifying exceptional motifs is often used for extracting information from long DNA sequences. The two difficulties of the method are the choice of the model that defines the expected frequencies of words and the approximation of the variance of the difference T(W) between the number of occurrences of a word W and its estimation. We consider here different Markov chain models, either with stationary or periodic transition probabilities. We estimate the variance of the difference T(W) by the conditional variance of the number of occurrences of W given the oligonucleotides counts that define the model. Two applications show how to use asymptotically standard normal statistics associated with the counts to describe a given sequence in terms of its outlying words. Sequences of Escherichia coli and of Bacillus subtilis are compared with respect to their exceptional tri- and tetranucleotides. For both bacteria, exceptional 3-words are mainly found in the coding frame. E. coli palindrome counts are analyzed in different models, showing that many overabundant words are one-letter mutations of avoided palindromes.

Bacillus subtilis↗

DNA sequence and analysis of human chromosome 18.

Chromosome 18 appears to have the lowest gene density of any human chromosome and is one of only three chromosomes for which trisomic individuals survive to term. There are also a number of genetic disorders stemming from chromosome 18 trisomy and aneuploidy. Here we report the finished sequence and gene annotation of human chromosome 18, which will allow a better understanding of the normal and disease biology of this chromosome. Despite the low density of protein-coding genes on chromosome 18, we find that the proportion of non-protein-coding sequences evolutionarily conserved among mammals is close to the genome-wide average. Extending this analysis to the entire human genome, we find that the density of conserved non-protein-coding sequences is largely uncorrelated with gene density. This has important implications for the nature and roles of non-protein-coding sequence elements.

Aneuploidy↗

Restricted and shared patterns of TCR beta-chain gene expression in silicone breast implant capsules and remote sites of tissue inflammation.

Silicone breast implants (SBI) induce formation of a periprosthetic, often inflammatory, fibrovascular neo-tissue called a capsule. Histopathology of explanted capsules varies from densely fibrotic, acellular specimens to those showing intense inflammation with activated macrophages, multinucleated giant cells, and lymphocytic infiltrates. It has been proposed that capsule-infiltrating lymphocytes comprise a secondary, bystander component of an otherwise benign foreign body response in women with SBIs. In symptomatic women with SBIs, however, the relationship of capsular inflammation to inflammation in other remote tissues remains unclear. In the present study, we utilized a combination of TCR beta-chain CDR3 spectratyping and DNA sequence analysis to assess the clonal heterogeneity of T cells infiltrating SBI capsules and remote, inflammatory tissues. TCR CDR3 fragment analysis of 22 distinct beta variable (BV) gene families revealed heterogeneous patterns of T cell infiltration in patients' capsules. In some cases, however, TCR BV transcripts exhibiting restricted clonality with shared CDR3 lengths were detected in left and right SBI capsules and other inflammatory tissues. DNA sequence analysis of shared, size-restricted CDR3 fragments confirmed that certain TCR BV transcripts isolated from left and right SBI capsules and multiple, extracapsular tissues had identical amino acid sequences within the CDR3 antigen binding domain. These data suggest that shared, antigen-driven T cell responses may contribute to chronic inflammation in SBI capsules as well as systemic sites of tissue injury.

Adult↗

Human chromosome 11 DNA sequence and analysis including novel gene identification.

Chromosome 11, although average in size, is one of the most gene- and disease-rich chromosomes in the human genome. Initial gene annotation indicates an average gene density of 11.6 genes per megabase, including 1,524 protein-coding genes, some of which were identified using novel methods, and 765 pseudogenes. One-quarter of the protein-coding genes shows overlap with other genes. Of the 856 olfactory receptor genes in the human genome, more than 40% are located in 28 single- and multi-gene clusters along this chromosome. Out of the 171 disorders currently attributed to the chromosome, 86 remain for which the underlying molecular basis is not yet known, including several mendelian traits, cancer and susceptibility loci. The high-quality data presented here--nearly 134.5 million base pairs representing 99.8% coverage of the euchromatic sequence--provide scientists with a solid foundation for understanding the genetic basis of these disorders and other biological phenomena.

Chromosomes, Human, Pair 11↗

DNA sequencing and analysis of the low-Ca2+-response plasmid pCD1 of Yersinia pestis KIM5.

The low-Ca2+-response (LCR) plasmid pCD1 of the plague agent Yersinia pestis KIM5 was sequenced and analyzed for its genetic structure. pCD1 (70,509 bp) has an IncFIIA-like replicon and a SopABC-like partition region. We have assigned 60 apparently intact open reading frames (ORFs) that are not contained within transposable elements. Of these, 47 are proven or possible members of the LCR, a major virulence property of human-pathogenic Yersinia spp., that had been identified previously in one or more of Y. pestis or the enteropathogenic yersiniae Yersinia enterocolitica and Yersinia pseudotuberculosis. Of these 47 LCR-related ORFs, 35 constitute a continuous LCR cluster. The other LCR-related ORFs are interspersed among three intact insertion sequence (IS) elements (IS100 and two new IS elements, IS1616 and IS1617) and numerous defective or partial transposable elements. Regional variations in percent GC content and among ORFs encoding effector proteins of the LCR are additional evidence of a complex history for this plasmid. Our analysis suggested the possible addition of a new Syc- and Yop-encoding operon to the LCR-related pCD1 genes and gave no support for the existence of YopL. YadA likely is not expressed, as was the case for Y. pestis EV76, and the gene for the lipoprotein YlpA found in Y. enterocolitica likely is a pseudogene in Y. pestis. The yopM gene is longer than previously thought (by a sequence encoding two leucine-rich repeats), the ORF upstream of ypkA-yopJ is discussed as a potential Syc gene, and a previously undescribed ORF downstream of yopE was identified as being potentially significant. Eight other ORFs not associated with IS elements were identified and deserve future investigation into their functions.

Calcium↗

The use of real-time PCR methods in DNA sequence variation analysis.

BACKGROUND: Real-time (RT) PCR methods for discovering and genotyping single nucleotide polymorphisms (SNPs) are becoming increasingly important in various fields of biological sciences. SNP genotyping is widely used to perform genetic association studies aimed at characterising the genetic factors underlying inherited traits. The detection and quantification of somatic mutations is an important tool for investigating the genetic causes of tumorigenesis. In infectious disease diagnostics there is an increasing emphasis placed on genotyping variation within the genomes of pathogenic organisms in order to distinguish between strains. METHODS: There are several platforms and methods available to the researcher wishing to undertake SNP analysis using real-time PCR methods. These use fluorescent technologies for discriminating between the alternate alleles of a polymorphism. There are several real-time PCR platforms currently on the market. Two of the key technical challenges are allele discrimination and allele quantification. CONCLUSIONS: Applications of this technology include SNP genotyping, the sensitive detection of somatic mutations and infectious disease subtyping.

Animals↗