Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Nucleotide Mapping”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Common sites of retroviral integration in mouse hematopoietic tumors identified by high-throughput, single nucleotide polymorphism-based mapping and bacterial artificial chromosome hybridization.

Retroviral insertional mutagenesis in mouse hematopoietic tumors provides a powerful cancer gene discovery tool. Here, we describe a high-throughput, single nucleotide polymorphism (SNP)-based method, for mapping retroviral integration sites cloned from mouse tumors, and a bacterial artificial chromosome (BAC) hybridization method, for localizing these retroviral integration sites to common sites of retroviral integration (CISs). Several new CISs were identified, including one CIS that mapped near Notch1, a gene that has been causally associated with human T-cell tumors. This mapping method is applicable to many different species, including ones where few genetic markers and little genomic sequence information are available. It can also be used to map endogenous proviruses.

Animals↗

Fractal dimension of exon and intron sequences.

In this paper, the concept of fractal is applied to describe the features of nucleotide sequences. We introduce the mapping from nucleotide sequences to two-dimensional metric space. Then we use this mapping to study quantitatively the self-similarity of exon and intron sequences in different scales. We find that self-similarity exists in the geometrical range and main range of a nucleotide sequence and define the fractal dimension in these ranges. The results show that the fractal properties of exon sequences are quite different from those of introns, reflecting their difference in structure and function. The fractal dimension of the geometrical range may be used to predict the exon regions of a raw nucleotide sequence.

Animals↗

Nucleotide sequence and genetic map of the 16-kb vaccinia virus HindIII D fragment.

We have determined the nucleotide sequence of the 16,059-bp HindIII D fragment from vaccinia virus strain WR. Translation in all 6 reading frames reveals a set of 22 open reading frames (ORFs), which are capable of encoding proteins ranging from 61 to 844 amino acids in length. With one exception, ORF 12, we have divided them into two primary sets according to their size. The minor group contains eight members ranging in length from 61 to 84 amino acids. The major group has thirteen members varying from 146 to 844 amino acids in length, and, in addition, due to its location on the DNA, one small ORF, 61 amino acids long. The neighboring major ORFs are closely packed along the DNA, being separated by 42 or fewer base pairs. In several instances the ends of adjoining ORFs overlap for up to 11 triplet codons. In three cases, 1 or 2 bases are shared between translation start and stop signals in adjacent ORFs. Regions of both strands of the DNA are transcribed. Two sets of temperature-sensitive mutations, totaling 17, which map to the HindIII D fragment, have been combined into eight complementation groups. The results of marker rescue analysis map one or more member of each group to a site in the HindIII D fragment within a defined open reading frame.

Amino Acid Sequence↗

Use of single nucleotide polymorphism-based mapping arrays to detect copy number changes and loss of heterozygosity in multiple myeloma.

The genetics of multiple myeloma is a vastly studied field in which techniques such as classical cytogenetics, fluorescence in situ hybridization, and comparative genomic hybridization have been used. More recently, single nucleotide polymorphism (SNP)-based mapping arrays have become available that allow the identification of regions of gain or loss as small as 2.5 kb. In addition to the increased resolution of SNP-based arrays, the detection of loss of heterozygosity is also possible. This allows the identification of loss of heterozygosity regions that arise through monosomy and recombination, resulting in uniparental disomy, which cannot be detected by conventional genetic methods. In this review, we discuss the benefits of SNP-based arrays along with some of the drawbacks and how that data can be used in conjunction with expression data to identify genes with altered expression in regions of interest.

Chromosome Mapping↗

Identification, cloning, nucleotide sequence and chromosomal map location of hns, the structural gene for Escherichia coli DNA-binding protein H-NS.

Beginning with a synthetic oligonucleotide probe derived from its amino acid sequence, we have identified, cloned and sequenced the hns gene encoding H-NS, an abundant Escherichia coli 15 kDa DNA-binding protein with a possible histone-like function. The amino acid sequence of the protein deduced from the nucleotide sequence is in full agreement with that determined for H-NS. By comparison of the restriction map of the cloned gene and of its neighboring regions with the physical map of E. coli K12 as well as by hybridization of the hns gene with restriction fragments derived from the total chromosome, we have located the hns gene oriented counterclockwise at 6.1 min on the E. coli chromosome, just before an IS30 insertion element.

Amino Acid Sequence↗

Detailed transcription map of Aleutian mink disease parvovirus.

We studied the transcription program of Aleutian mink disease parvovirus (ADV) by using a combination of cDNA cloning and sequencing, primer extension, and Northern (RNA) blot hybridization with splice-specific oligonucleotides. The 4.8-kilobase ADV genome was transcribed in the rightward direction, yielding plus-sense polyadenylated transcripts of 4.3 (R1 RNA), 2.8 (R2), 2.8 (R3), 1.1 (RX), and 0.85 (R2') kilobases. Each RNA transcript had potential translation initiation sites within open reading frames, suggesting protein translation, and a scheme encompassing ADV structural and nonstructural proteins is proposed. Each of the five RNA transcripts had a characteristic set of splices and originated from a promoter at nucleotide 152 (map unit 3 [R1, R2, R2', and RX]) or at nucleotide 1729 (map unit 36 [R3]). The transcripts terminated with a poly(A) tail at one of two positions: either at map unit 53 (R2' and RX) or at map unit 92 (R1, R2, and R3). Similarities with and differences from the transcription maps of other parvoviruses are discussed, and possible roles of the unique features found in ADV transcription are related to the special pathogenic features of this virus.

Aleutian Mink Disease Virus↗

Nucleotide sequence and transcript map of the Agrobacterium tumefaciens Ti plasmid-encoded octopine synthase gene.

We have determined the complete nucleotide sequence of the gene for the crown gall enzyme, octopine synthase. The sequence was derived from cloned fragments of the Agrobacterium tumefaciens Ti plasmid Ach5. It displayed a continuous open reading frame encoding a polypeptide chain of 358 amino acids. The nucleotide positions corresponding to the 5' end and poly(A) addition site of the mature octopine synthase mRNA from a tobacco tumor cell line were determined by S1 nuclease mapping. Two sequences closely resembling transcriptional control regions found in eukaryotic genes transcribed by RNA polymerase II were identified in the flanking genomic DNA: a sequence 5'-TATTTAAA-3' was located 32 base pairs upstream from the initiation site of transcription, and a hexanucleotide 5'-AATAAT-3' occurred 17 base pairs in front of the poly(A) addition site. No Shine-Dalgarno sequence was present in the untranslated 5' leader sequence. The observations indicate that this DNA sequence, although naturally carried by a bacterial plasmid, is programmed as a functional plant gene.

Amino Acid Oxidoreductases↗

Single nucleotide polymorphisms (SNPs) that map to gaps in the human SNP map.

An international effort is underway to generate a comprehensive haplotype map (HapMap) of the human genome represented by an estimated 300,000 to 1 million 'tag' single nucleotide polymorphisms (SNPs). Our analysis indicates that the current human SNP map is not sufficiently dense to support the HapMap project. For example, 24.6% of the genome currently lacks SNPs at the minimal density and spacing that would be required to construct even a conservative tag SNP map containing 300,000 SNPs. In an effort to improve the human SNP map, we identified 140,696 additional SNP candidates using a new bioinformatics pipeline. Over 51,000 of these SNPs mapped to the largest gaps in the human SNP map, leading to significant improvements in these regions. Our SNPs will be immediately useful for the HapMap project, and will allow for the inclusion of many additional genomic intervals in the final HapMap. Nevertheless, our results also indicate that additional SNP discovery projects will be required both to define the haplotype architecture of the human genome and to construct comprehensive tag SNP maps that will be useful for genetic linkage studies in humans.

Base Sequence↗

Nucleotide sequence and linkage map position of the genes for ribosomal proteins L14 and S8 in the maize chloroplast genome.

The nucleotide sequence of a 1287-base-pair segment of the maize (Zea mays) chloroplast DNA, encoding chloroplast ribosomal proteins L14, S8 and the C-terminal part of L16, has been determined using the dideoxy-chain-termination method. These data from a monocot plant are compared to the corresponding data from a dicot and a lower plant and from two bacteria. The deduced amino acid sequence of maize chloroplast L14 shows 80%, 81%, 51% and 52% and that of S8 shows 75%, 58%, 39% and 38% sequence identity, respectively, to the corresponding sequences of Nicotiana tabacum, Marchantia polymorpha, Bacillus stearothermophilus and Escherichia coli. The starting map coordinates of rpL14 and rpS8 in the physical map of the maize chloroplast DNA [Larrinua, I. M., Muskavitch, K. M. T., Gubbins, E. J. and Bogorad, L. (1983) Plant Mol. Biol. 2, 129-140] are 31.330 and 31.841. The gene order is rpL16-spacer-rpL14-spacer-rpS8. Shine-Dalgarno sequences (GGA and AGGAGG) and computer-derived stem-loop structures of dyad symmetry are present in the spacers and the 3' downstream region of rpS8, respectively, but a chloroplast promoter-like sequence could not be detected suggesting that the latter might be located further upstream in this ribosomal protein gene cluster in maize chloroplast DNA.

Amino Acid Sequence↗

Nucleotide sequence and linkage map position of the secX gene in maize chloroplast and evidence that it encodes a protein belonging to the 50S ribosomal subunit.

The nucleotide sequence of the segment of maize chloroplast DNA lying between the map coordinate positions 32.59 and 32.98 Kb and containing the secX gene has been determined. The derived amino acid sequence of maize chloroplast secX is 95%, 87% and 62% identical to the corresponding derived amino acid sequences from two plant chloroplasts and Escherichia coli, respectively. It is also 70% identical to the experimentally determined amino acid sequence of a protein isolated from Bacillus stearothermophilus ribosomes. Separation of the 50S ribosomal subunit proteins of E. coli by reversed phase HPLC gave a peak which contained pure secX protein, as determined by N-terminal amino acid sequencing. Spinach chloroplast 50S subunit proteins separated by HPLC also gave a peak corresponding to pure secX protein. From these results we conclude that the secX gene in E. coli and in plant chloroplasts encodes a small (37-38 amino acid residues) ribosomal protein belonging to the 50S subunit. The same conclusion has been reached recently by A. Wada with respect to E. coli secX. In agreement with Wada, we name the secX protein L36. Its chloroplast gene is designated rpL36.

Base Sequence↗

Localization of cancer susceptibility genes by genome-wide single-nucleotide polymorphism linkage-disequilibrium mapping.

With the large numbers of single nucleotide polymorphisms (SNPs) available and new technologies that permit high throughput genotyping, we have investigated the possibility of the localization of disease genes with genome-wide panels of SNP markers and taking advantage of the linkage-disequilibrium (LD) between the disease gene and closely linked markers. For this purpose, we selected cases from the Ashkenazi Jewish population, in which the mutant alleles are expected to be identical by descent from a common founder and the regions of LD encompassing these mutant alleles are large. As a validation of this approach for localization, we performed two trials: one in autosomal recessive Bloom syndrome, in which a unique mutation of the BLM gene is present at elevated frequencies in cases, and the other in autosomal dominant hereditary nonpolyposis colorectal cancer (HNPCC), in which a unique mutation of MSH2 is present at elevated frequencies. In the Bloom syndrome trial, we genotyped 3,258 SNPs in 10 Jewish Bloom syndrome cases and 31 non-Bloom syndrome Jewish persons as a comparison group. In the HNPCC trial, we genotyped 8,549 SNPS in 13 Jewish HNPCC cases whose colon cancers exhibited microsatellite instability and in 63 healthy Jews as a comparison group. To identify significant associations, we performed (a) Fisher's exact test comparing genotypes at each locus in cases versus controls and (b) a haplotype analysis by estimating the frequency of haplotypes with the expectation-maximization algorithm and comparing haplotype frequencies in cases versus controls by logistic regression and a maximum likelihood ratio method. In the Bloom syndrome trial, by Fisher's exact test, statistically significant association was detected at a single locus, TSC0754862, which is a locus 1.7 million bp from BLM. Two-locus, three-locus, and four-locus haplotypes that included TSC0754862 and flanked BLM were also statistically more frequent in cases versus controls. In the HNPCC trial, although a significant P value was not obtained by the single SNP genotype analysis, significant associations were detected for several multilocus haplotypes in an 11-million-bp region that contained the MSH2 gene. This work demonstrates the power of the LD mapping approach in an isolated population and its general applicability to the identification of novel cancer-causing genes.

Adenosine Triphosphatases↗

Nucleotide sequence and promoter mapping of the Escherichia coli Shiga-like toxin operon of bacteriophage H-19B.

We determined the nucleotide sequence of the Shiga-like toxin-1 (SLT-1) genes carried by the toxin-converting bacteriophage H-19B. Two open reading frames were identified; these were separated by 12 base pairs and encoded proteins of 315 (A subunit) and 89 (B subunit) amino acids. The predicted protein subunits had N-terminal hydrophobic signal sequences of 22 and 20 amino acids, respectively. The predicted amino acid sequence of the B subunit was identical to that of the B subunit of Shiga toxin. The A chain of ricin was found to be significantly related to the predicted A1 fragment of the SLT-1 A subunit. S1 nuclease protection experiments showed that the two cistrons formed a single transcriptional unit, with the A subunit being proximal to the promoter. A probable promoter was identified by primer extension, and transcription was found to increase dramatically under conditions of iron starvation. A 21-base-pair sequence with dyad symmetry was found in the region of the SLT-1 -10 sequence, which was found to be 68% homologous to a region of dyad symmetry found in the -35 region of the promoter of the iucA gene on plasmid ColV-K30, which specifies the 74,000-dalton ferric-aerobactin receptor protein. Betley et al. (M. Betley, V. Miller, and J. Mekalanos, Annu. Rev. Microbiol. 40:577-605, 1986) have recently summarized evidence suggesting that the slt operon is under the control of the fur regulatory system. The area of dyad symmetry found in both promoters may represent a regulatory site. A rho-independent terminator sequence was found 230 base pairs downstream from the B cistron stop codon.

Amino Acid Sequence↗

Nucleotide sequence and genetic map of cowpea severe mosaic virus RNA 2 and comparisons with RNA 2 of other comoviruses.

We report the nucleotide sequence of cowpea severe mosaic comovirus (CPSMV) genomic RNA 2. The molecule is composed of 3732 nucleotide (nt) residues, exclusive of the polyadenylate at the 3' end. Only one of the six reading frame registers has a long open reading frame, from nt 255 to nt 3260 in the polarity of encapsidated RNA and corresponding to a polyprotein of 1002 amino acid residues (aa). As has been reported for other comoviruses, a second in-frame AUG, at nt position 531, apparently also initiates translation, at least in vitro. Multiple alignments of the deduced CPSMV polyprotein aa sequence with those of bean pod mottle comovirus (BPMV), cowpea mosaic comovirus (CPMV), and red clover mottle comovirus (RCMV) were consistent with a similar size for each of the three genes: the putative movement protein, beginning at the second in-frame AUG, the large coat protein (L), and the small coat protein. Identical nucleotide sequences in the terminal noncoding regions of RNA 2 of the four viruses are limited to 9 nt at the 5' end and the 3' polyadenylate. However, extensive similarities in sequence and potential structure were found. For all three genes and the 5' untranslated region, CPSMV and BPMV are more similar to each other than either is to CPMV or RCMV, the last two being similar to each other. Observed similarities predict that both cleavage sites in the CPSMV RNA 2 polyprotein are at glutamine-serine dipeptides. A sequence of 16 aa at the amino terminus of L, determined by automated Edman degradation, matched a region of the deduced aa sequence in the polyprotein and is consistent with cleavage at the predicted glutamine-serine dipeptide.

Amino Acid Sequence↗

Nucleotide sequence and transcript mapping of the tmr gene of the pTiA6NC octopine Ti-plasmid: a bacterial gene involved in plant tumorigenesis.

The nucleotide sequence of a tumor morphology gene, tmr, from the Agrobacterium tumefaciens Ti-plasmid, pTiA6NC, and its flanking 5' region was determined by M13 "dideoxy" procedures. The DNA sequence reveals an open reading frame capable of encoding a 240 amino acid protein. We have identified the polyadenylated transcript initiation and termination sits by S1 nuclease mapping. The extent of the sequence required for transcription 5' to the start of transcription has been delimited by two transposon insertions. The first of these maps at -- 121 with respect to transcription initiation and results in the wild-type phenotype, the second insertion maps at about -85 and results in a tmr phenotype.

Arginine↗

Quantitative trait loci mapped to single-nucleotide resolution in yeast.

Identifying the genetic variation underlying quantitative trait loci remains problematic. Consequently, our molecular understanding of genetically complex, quantitative traits is limited. To address this issue directly, we mapped three quantitative trait loci that control yeast sporulation efficiency to single-nucleotide resolution in a noncoding regulatory region (RME1) and to two missense mutations (TAO3 and MKT1). For each quantitative trait locus, the responsible polymorphism is rare among a diverse set of 13 yeast strains, suggestive of genetic heterogeneity in the control of yeast sporulation. Additionally, under optimal conditions, we reconstituted approximately 92% of the sporulation efficiency difference between the two genetically distinct parents by engineering three nucleotide changes in the appropriate yeast genome. Our results provide the highest resolution to date of the molecular basis of a quantitative trait, showing that the interaction of a few genetic variants can have a profound phenotypic effect.

Adaptor Proteins, Signal Transducing↗

Nucleotide sequence and transcript mapping of the HindIII F region of the Autographa californica nuclear polyhedrosis virus genome.

The organization of genes in the 2.95 kb EcoRI-SalI fragment (0.7 to 3.0 map units) located within the HindIII F region of the Autographa californica nuclear polyhedrosis virus genome was studied by a combination of DNA sequencing, Northern blot analysis, S1 mapping and primer extension analysis. In addition to the two divergent overlapping transcripts [leftward early (ES1) and rightward late (ES2)] previously reported, a third transcript which is present at 18 h p.i. and which also runs leftward, overlapping ES1 by 1600 nucleotides (nt) at the 3' end, was mapped to this region. The DNA sequence revealed the presence of three open reading frames (ORFs) of significant length. ORF-1 and ORF-2 correspond to the leftward transcripts, and code for potential polypeptides of 151 and 329 amino acids, respectively. ORF-3 which codes for a potential polypeptide of 167 amino acids is located on the opposite strand in a region for which no transcript mapping data are available. However, the conserved late gene promoter/cap site sequence (ATAAG) is present 23 nt upstream of the start of ORF-3.

Amino Acid Sequence↗

Identification of a stable RNA encoded by the H-strand of the mouse mitochondrial D-loop region and a conserved sequence motif immediately upstream of its polyadenylation site.

By using a combination of Northern blot hybridization with strand-specific DNA probes, S1 nuclease protection, and sequencing of oligo-dT-primed cDNA clones, we have identified a 0.8 kb poly(A)-containing RNA encoded by the H-strand of the mouse mitochondrial D-loop region. The 5' end of the RNA maps to nucleotide 15417, a region complementary to the start of tRNA(Pro) gene and the 3' polyadenylated end maps to nucleotide 16295 of the genome, immediately upstream of tRNA(Phe) gene. The H-strand D-loop region encoded transcripts of similar size are also detected in other vertebrate systems. In the mouse, rat, and human systems, the 3' ends of the D-loop encoded RNA are preceded by conserved sequences AAUAAA, AAUUAA, or AACUAA, that resemble the polyadenylation signal. The steady-state level of the RNA is generally low in dividing or in vitro cultured cells, and markedly higher in differentiated tissues like liver, kidney, heart, and brain. Furthermore, an over 10-fold increase in the level of this RNA is observed during the induced differentiation of C2C12 mouse myoblast cells into myotubes. These results suggest that the D-loop H-strand encoded RNA may have yet unknown biological functions. A 20 base pair DNA sequence from the 3' terminal region containing the conserved sequence motif binds to a protein from the mitochondrial extracts in a sequence-specific manner. The binding specificity of this protein is distinctly different from the previously characterized H-strand DNA termination sequence in the D-loop or the H-strand transcription terminator immediately downstream of the 16S rRNA gene. Thus, we have characterized a novel poly(A)-containing RNA encoded by the H-strand of the mitochondrial D-loop region and also identified the putative ultimate termination site for the H-strand transcription.

Adenine↗

Little loss of information due to unknown phase for fine-scale linkage-disequilibrium mapping with single-nucleotide-polymorphism genotype data.

We present the results of a simulation study that indicate that true haplotypes at multiple, tightly linked loci often provide little extra information for linkage-disequilibrium fine mapping, compared with the information provided by corresponding genotypes, provided that an appropriate statistical analysis method is used. In contrast, a two-stage approach to analyzing genotype data, in which haplotypes are inferred and then analyzed as if they were true haplotypes, can lead to a substantial loss of information. The study uses our COLDMAP software for fine mapping, which implements a Markov chain-Monte Carlo algorithm that is based on the shattered coalescent model of genetic heterogeneity at a disease locus. We applied COLDMAP to 100 replicate data sets simulated under each of 18 disease models. Each data set consists of haplotype pairs (diplotypes) for 20 SNPs typed at equal 50-kb intervals in a 950-kb candidate region that includes a single disease locus located at random. The data sets were analyzed in three formats: (1). as true haplotypes; (2). as haplotypes inferred from genotypes using an expectation-maximization algorithm; and (3). as unphased genotypes. On average, true haplotypes gave a 6% gain in efficiency compared with the unphased genotypes, whereas inferring haplotypes from genotypes led to a 20% loss of efficiency, where efficiency is defined in terms of root mean integrated square error of the location of the disease locus. Furthermore, treating inferred haplotypes as if they were true haplotypes leads to considerable overconfidence in estimates, with nominal 50% credibility intervals achieving, on average, only 19% coverage. We conclude that (1). given appropriate statistical analyses, the costs of directly measuring haplotypes will rarely be justified by a gain in the efficiency of fine mapping and that (2). a two-stage approach of inferring haplotypes followed by a haplotype-based analysis can be very inefficient for fine mapping, compared with an analysis based directly on the genotypes.

Algorithms↗