Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Contig Mapping”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,531 records · Page 85Linked to original sources

Fine mapping of the cystinosis gene using an integrated genetic and physical map of a region within human chromosome band 17p13.

The cystinosis gene has been reported to reside in a 3.1 cM region of chromosome 17p13 flanked by markers D17S1828 and D17S1798. We created a yeast artificial chromosome (YAC) contig between these markers and report here an integrated genetic and physical map which will aid in the identification of other genes in this area. Using one pertinent YAC clone, 898A10, we identified new polymorphic markers in the cystinosis gene region. One such marker, D17S2167, was localized by radiation hybrid analysis to within 10.2 cR8000 of D17S1828. Haplotype analysis in two separate informative families revealed recombination events which placed the cystinosis gene between markers D17S1828 and D17S2167, an area estimated to be 187-510 kb in size. This dramatic narrowing of the cystinosis gene region permits the creation of a P1 or cosmid contig across the area of interest. The ultimate cloning of the cystinosis gene should eventually reveal how a functional lysosomal transport protein is synthesized, targetted, processed, and integrated into the lysosomal membrane.

Base Sequence↗

Isolation of a human YAC contig encompassing a cluster of UGT2 genes and its regional localization to chromosome 4q13.

Previously we mapped the gene encoding a human bile acid UDP-glucuronosyltransferase (UGT2B4) to chromosome 4. Here we report the mapping of two additional human UGT2B genes to chromosome 4 utilizing the polymerase chain reaction (PCR) and a panel of human/rodent somatic cell hybrid cell lines. A yeast artificial chromosome contig containing the UGT2B4, UGT2B9, and UGT2B15 genes was isolated, and pulsed-field gel electrophoresis and PCR revealed that several members of the human UGT2B gene subfamily are clustered within a 195-kb region of the YAC contig. These data permitted a provisional ordering of the genes as UGT2B9-UGT2B4-UGT2B15. Fluorescence in situ hybridization analysis, using the YAC DNA, permitted the regional localization of this gene cluster to chromosome 4q13.

Animals↗

Multiple-complete-digest restriction fragment mapping: generating sequence-ready maps for large-scale DNA sequencing.

Multiple-complete-digest mapping is a DNA mapping technique based on complete-restriction-digest fingerprints of a set of clones that provides highly redundant coverage of the mapping target. The maps assembled from these fingerprints order both the clones and the restriction fragments. Maps are coordinated across three enzymes in the examples presented. Starting with yeast artificial chromosome contigs from the 7q31.3 and 7p14 regions of the human genome, we have produced cosmid-based maps spanning more than one million base pairs. Each yeast artificial chromosome is first subcloned into cosmids at a redundancy of x15-30. Complete-digest fragments are electrophoresed on agarose gels, poststained, and imaged on a fluorescent scanner. Aberrant clones that are not representative of the underlying genome are rejected in the map construction process. Almost every restriction fragment is ordered, allowing selection of minimal tiling paths with clone-to-clone overlaps of only a few thousand base pairs. These maps demonstrate the practicality of applying the experimental and software-based steps in multiple-complete-digest mapping to a target of significant size and complexity. We present evidence that the maps are sufficiently accurate to validate both the clones selected for sequencing and the sequence assemblies obtained once these clones have been sequenced by a "shotgun" method.

Base Composition↗

Physical mapping and cloning of the proximal segment of the myotonic dystrophy gene region.

The myotonic dystrophy (DM) region has been recently shown to be bracketed by two key recombinant events. One recombinant occurs in a Dutch DM family, which maps the DM locus distal to the ERCC1 gene and D19S115 (pE0.8). The other recombinant event is in a French Canadian DM family, which maps DM proximal to D19S51 (p134c). To further resolve this region, we initiated a chromosome walk in a telomeric direction from pE0.8, a proximal marker tightly linked to DM, toward the genetic locus. An Alu-PCR approach to chromosome walking in a cosmid library from flow-sorted chromosome 19 was used to isolate DM region cosmids. This effort has resulted in the cloning of a 350-kb genomic contig of human chromosome 19q13.3. New genetic and physical mapping information has been generated using the newly cloned markers from this study. As a result of this new mapping information, the minimal area that is to contain the DM gene has been redefined. Approximately 200 kb of sequence between pE0.8 and the closest proximal marker to DM, pKEX0.8, that would have otherwise been screened for DM candidate genes, has been eliminated as containing the DM gene.

Base Sequence↗

A genetic interval and physical contig spanning the Peronospora parasitica (At) avirulence gene locus ATR1Nd.

In Peronospora parasitica (At) (downy mildew), the genetic determinants of cultivar-specific recognition by Arabidopsis thaliana are the Arabidopsis thaliana-recognised (ATR) avirulence genes. We describe the identification of 10 amplified fragment length polymorphism (AFLP) markers that define a genetic mapping interval for the ATR1Nd avirulence allele, the presence of which is perceived by the RPP1Nd resistance gene. Furthermore, we have constructed a P. parasitica (At) bacterial artificial chromosome (BAC) library comprising over 630Mb of cloned DNA. We have isolated 16 overlapping clones from the BAC library that form a contig spanning the genetic interval. BAC sequence-derived markers and a total mapping population of 311 F(2) individuals were used to refine the ATR1Nd locus to a 1cM interval that is represented by four BAC clones and spans less than 250kb of DNA. This work demonstrates that map-based cloning techniques are feasible in this organism and provides the critical foundations for cloning ATR1Nd using such a strategy.

Alleles↗

Cosmid contigs spanning 9q34 including the candidate region for TSC1.

The tuberous sclerosis disease gene TSC1 has been mapped to 9q34. However, its precise localisation has proved problematic because of conflicting recombination data. Therefore, we have attempted to clone the entire target area into cosmid contigs prior to gene isolation studies. We have used Alu-PCR from irradiation hybrids to produce complex probes from the target region which have identified 1,400 cosmids from a chromosome-specific library. These, along with cosmids obtained by other methods, have been assembled into contigs by a fingerprinting technique. We estimate that we have obtained most of the region in cosmid contigs. These cosmids are a resource for the isolation of expressed genes within the TSC1 interval. In addition, the cosmid contig assembly has demonstrated a number of previously unknown physical connections between genes and markers in 9q34.

Animals↗

IRS-Bubble PCR: An Effective Method for Representative Amplification of Human Genomic DNA Sequences from Complex Sources

A rate-limiting step in the analysis of large segments of genomic DNA is the generation of a set of representative short single-copy sequences that can be used for development of fine-structure maps. In direct response to this need, we have developed interspersed repetitive DNA sequence (IRS)-bubble PCR. IRS-bubble PCR was designed to amplify the human DNA content of somatic cell hybrids, yeast artificial chromosomes (YACs), BACs, PACs, cosmids, and lambda phage and to result in greater complexity and representation than standard inter-IRS PCR. Here, we describe the application of IRS-bubble PCR to the generation of complex clone libraries for targeted STS development and the construction of robust hybridization probes for FISH mapping, "chromosome painting," or the assembly of cosmid contigs representing the human DNA content of somatic cell hybrids or YACs.

Journal Article↗

Identification, characterisation and clinical applications of cosmids from the telomeric and centromeric regions of the long arm of chromosome 22.

Using human telomeric repeats and centromeric alpha repeats, we have identified adjacent single copy cosmid clones from human chromosome 22 cosmid libraries. These single copy cosmids were mapped to chromosome 22 by fluorescence in situ hybridisation (FISH). Based on these cosmids, we established contigs that included part of the telomeric and subtelomeric regions, and part of the centromeric and pericentromeric regions of the long arm of human chromosome 22. Each of the two cosmid contigs consisted of five consecutive steps and spanned approximately 100-150 kb at both extreme ends of 22q. Moreover, highly informative polymorphic markers were identified in the telomeric region. Our results suggest that the telomere specific repeat (TTAGGG)n encompasses a region that is larger than 40 kb. The cosmid contigs and restriction fragment length polymorphism markers described here are useful tools for physical and genetic mapping of chromosome 22, and constitute the basis of further studies of the structure of the subtelomeric and pericentromeric regions of 22q. We also demonstrate the use of these clones in clinical diagnosis of different chromosome 22 aberrations by FISH.

Base Sequence↗

Annotation and BAC/PAC localization of nonredundant ESTs from drought-stressed seedlings of an indica rice.

To decipher the genes associated with drought stress response and to identify novel genes in rice, we utilized 1540 high-quality expressed sequence tags (ESTs) for functional annotation and mapping to rice genomic sequences. These ESTs were generated earlier by 3'-end single-pass sequencing of 2000 cDNA clones from normalized cDNA libraries constructed form drought-stressed seedlings of an indica rice. A rice UniGene set of 1025 transcripts was constructed from this collection through the BLASTN algorithm. Putative functions of 559 nonredundant ESTs were identified by BLAST similarity search against public databases. Putative functions were assigned at a stringency E value of 10(-6) in BLASTN and BLASTX algorithms. To understand the gene structure and function further, we have utilized the publicly available finished and unfinished rice BAC/PAC (BAC, bacterial artificial chromosome; PAC, P1 artificial chromosome) sequences for similarity search using the BLASTN algorithm. Further, 603 nonredundant ESTs have been mapped to BAC/PAC clones. BAC clones were assigned by a homology of above 95% identity along 90% of EST sequence length in the aligned region. In all, 700 ESTs showed rice EST hits in GenBank. Of the 325 novel ESTs, 128 were localized to BAC clones. In addition, 127 ESTs with identified putative functions but with no homology in IRGSP (International Rice Genome Sequencing Program) BAC/PAC sequences were mapped to the Chinese WGS (whole genome shotgun contigs) draft sequence of the rice genome. Functional annotation uncovered about a hundred candidate ESTs associated with abiotic stress in rice and Arabidopsis that were previously reported based on microarray analysis and other studies. This study is a major effort in identifying genes associated with drought stress response and will serve as a resource to rice geneticists and molecular biologists.

Chromosomes, Artificial, Bacterial↗

Construction of two BAC libraries from the wild Mexican diploid potato, Solanum pinnatisectum, and the identification of clones near the late blight and Colorado potato beetle resistance loci.

To facilitate isolation and characterization of disease and insect resistance genes important to potato, two bacterial artificial chromosome (BAC) libraries were constructed from genomic DNA of the Mexican wild diploid species, Solanum pinnatisectum, which carries high levels of resistance to the most important potato pathogen and pest, the late blight and the Colorado potato beetle (CPB). One of the libraries was constructed from the DNA, partially digested with BamHI, and it consists of 40328 clones with an average insert size of 125 kb. The other library was constructed from the DNA partially digested with EcoRI, and it consists of 17280 clones with an average insert size of 135 kb. The two libraries, together, represent approximately six equivalents of the wild potato haploid genome. Both libraries were evaluated for contamination with organellar DNA sequences and were shown to have a very low percentage (0.65-0.91%) of clones derived from the chloroplast genome. High-density filters, prepared from the two libraries, were screened with ten restriction fragment length polymorphism (RFLP) markers linked to the resistance genes for late blight, CPB, Verticillium wilt and potato cyst nematodes, and the gene Sr1 for the self-incompatibility S-locus. Thirty nine positive clones were identified and at least two positive BAC clones were detected for each RFLP marker. Four markers that are linked to the late blight resistance gene Rpi1 hybridized to 14 BAC clones. Fifteen BAC clones were shown to harbor the PPO (polyphenol oxidase) locus for the CPB resistance by three RFLP probes. Two RFLP markers detected five BAC clones that were linked to the Sr1 gene for self-incompatibility. These results agree with the library's predicted extent of coverage of the potato genome, and indicated that the libraries are useful resources for the molecular isolation of disease and insect resistance genes, as well as other economically important genes in the wild potato species. The development of the two potato BAC libraries provides a starting point, and landmarks for BAC contig construction and chromosome walking towards the map-based cloning of agronomically important target genes in the species.

Animals↗

Elucidation of the exon-intron structure and size of the human protein kinase C beta gene (PRKCB).

As part of a transcriptional mapping project on human chromosome 16p12, a genomic contig was constructed that spanned the alternatively spliced human protein kinase C beta gene (PRKCB). PRKCB was determined to consist of 18 exons covering approximately 375 kb, with a particularly large intron of over 150 kb between exons 2 and 3. PRKCB is nearly 19 times larger than the highly homologous Drosophila melanogaster protein kinase C gene (dPKC), which has a similar-sized open reading frame but only 13 exons. This increase in size has occurred mostly as a result of expansion of introns, with intron size in the human gene averaging 22 kb compared with 1.5 kb in dPKC. The difference in gene size correlates with the difference in genome size, with the human haploid genome being nearly 18 times larger than the 170 Mb Drosophila haploid genome.

Blotting, Southern↗

Metabolic highways of Neurospora crassa revisited.

This chapter describes the metabolic pathways for Neurospora crassa in the biosynthesis of amino acids, purines, pyrimidines, vitamins, and cofactors, and for glycolysis, the TCA and glyoxylate cycles and the initial stages of the pentose phosphate pathway. For each step in metabolism, the gene or genes within the genome sequence of the species is identified, correlations are made with previously identified genes, and new gene designations are assigned to others. For each gene, details given are the function of the gene product, contig location, comparison of the genetic and physical map location, Saccharomyces cerevisiae homolog, and perhaps others, and the level of similarity.

Amino Acids↗

Exon/intron structure of the human AF-4 gene, a member of the AF-4/LAF-4/FMR-2 gene family coding for a nuclear protein with structural alterations in acute leukaemia.

The AF-4 gene on human chromosome 4q21 is involved in reciprocal translocations to the ALL-1 gene on chromosome 11q23, which are associated with acute lymphoblastic leukaemias. A set of recombinant phage carrying genomic fragments for the coding region and flanking sequences of the AF-4 gene were isolated. Phage inserts were assembled into four contigs with 21 exons, and an intron phase map was produced enabling the interpretation of translocation-generated fusion proteins. The gene contains two alternative first exons, 1a and 1b, both including a translation initiation codon. The translocation breakpoint cluster region is flanked by exons 3 and 6 and two different polyadenylation signals were identified. Polyclonal antisera directed against three different portions of the AF-4 protein were produced and used to detect a 116 kD protein in cellular extracts of human B-lymphoblastoid and proB cell lines. In mitogen-stimulated human peripheral blood mononuclear cells the AF-4 antigen was predominantly located in the nucleus. The AF-4 gene is a member of the AF-4, LAF-4 and FMR-2 gene family. The members of this family encode serine-proline-rich proteins with properties of nuclear transcription factors. Comparison of AF-4 protein coding sequences with the LAF-4 and FMR-2 sequences revealed five highly conserved domains of potential functional relevance.

Amino Acid Sequence↗

AdDeam: a fast and scalable tool for estimating and clustering reference-level damage profiles.

MOTIVATION: DNA damage patterns, such as increased frequencies of C→T and G→A substitutions at fragment ends, are widely used in ancient DNA studies to assess authenticity and detect contamination. In metagenomic studies, fragments can be mapped against multiple references or de novo assembled contigs to identify those likely to be ancient. Generating and comparing damage profiles, however, can be both tedious and time-consuming. Although tools exist for estimating damage in single reference genomes and metagenomic datasets, none efficiently cluster damage patterns. RESULTS: To address this methodological gap, we developed AdDeam, a tool that combines rapid damage estimation with clustering for streamlined analyses and easy identification of potential contaminants or outliers. Our tool takes aligned ancient DNA (aDNA) fragments from various samples or contigs as input, computes damage patterns, clusters them, and outputs representative damage profiles per cluster, a probability of each sample pertaining to a cluster, as well as a Principal Component Analysis of the damage patterns for each sample for fast visualisation. We evaluated AdDeam on both simulated and empirical datasets. AdDeam effectively distinguishes different damage levels, such as uracil-DNA glycosylase-treated samples, sample-specific damages from specimens of different time periods, and can also distinguish between contigs containing modern or ancient fragments, providing a clear framework for aDNA authentication and facilitating large-scale analyses. AVAILABILITY AND IMPLEMENTATION: AdDeam is publicly available at https://github.com/LouisPwr/AdDeam and can also be installed via Bioconda. It is implemented in Python and C++. All analysis scripts and datasets are available at https://github.com/LouisPwr/AdDeamAnalysis and on Zenodo under: 10.5281/zenodo.15052427.

Software↗

A yeast artificial chromosome contig spanning the Charcot-Marie-Tooth disease type 1A duplication region.

A contiguous set of 43 overlapping yeast artificial chromosome (YAC) clones has been developed for the Charcot-Marie-Tooth disease type 1A (CMT1A) duplication region of chromosome 17p11.2. The contig spans approximately 2.0 Mb and can be represented in a minimum of five overlapping YACs. The YAC clones were isolated from two total human genomic YAC libraries and from YAC libraries made from rodent-human hybrid cell lines. YAC clones were isolated from the libraries by polymerase chain reaction (PCR) technique. Localization to chromosome 17p11.2 was confirmed by fluorescence in situ hybridization. Overlap between the YAC clones was detected by inter-Alu PCR amplification of the YACs and by cross hybridization of the YACs with YAC insert ends obtained by Vectorette PCR. This YAC contig is a useful resource for analyzing and mapping all the genes contained within the CMT1A duplication.

Base Sequence↗

High-resolution radiation hybrid map of wheat chromosome 1D.

Physical mapping methods that do not rely on meiotic recombination are necessary for complex polyploid genomes such as wheat (Triticum aestivum L.). This need is due to the uneven distribution of recombination and significant variation in genetic to physical distance ratios. One method that has proven valuable in a number of nonplant and plant systems is radiation hybrid (RH) mapping. This work presents, for the first time, a high-resolution radiation hybrid map of wheat chromosome 1D (D genome) in a tetraploid durum wheat (T. turgidum L., AB genomes) background. An RH panel of 87 lines was used to map 378 molecular markers, which detected 2312 chromosome breaks. The total map distance ranged from approximately 3,341 cR(35,000) for five major linkage groups to 11,773 cR(35,000) for a comprehensive map. The mapping resolution was estimated to be approximately 199 kb/break and provided the starting point for BAC contig alignment. To date, this is the highest resolution that has been obtained by plant RH mapping and serves as a first step for the development of RH resources in wheat.

Chromosome Breakage↗

The 630-kb lung cancer homozygous deletion region on human chromosome 3p21.3: identification and evaluation of the resident candidate tumor suppressor genes. The International Lung Cancer Chromosome 3p21.3 Tumor Suppressor Gene Consortium.

We used overlapping and nested homozygous deletions, contig building, genomic sequencing, and physical and transcript mapping to further define a approximately 630-kb lung cancer homozygous deletion region harboring one or more tumor suppressor genes (TSGs) on chromosome 3p21.3. This location was identified through somatic genetic mapping in tumors, cancer cell lines, and premalignant lesions of the lung and breast, including the discovery of several homozygous deletions. The combination of molecular manual methods and computational predictions permitted us to detect, isolate, characterize, and annotate a set of 25 genes that likely constitute the complete set of protein-coding genes residing in this approximately 630-kb sequence. A subset of 19 of these genes was found within the deleted overlap region of approximately 370-kb. This region was further subdivided by a nesting 200-kb breast cancer homozygous deletion into two gene sets: 8 genes lying in the proximal approximately 120-kb segment and 11 genes lying in the distal approximately 250-kb segment. These 19 genes were analyzed extensively by computational methods and were tested by manual methods for loss of expression and mutations in lung cancers to identify candidate TSGs from within this group. Four genes showed loss-of-expression or reduced mRNA levels in non-small cell lung cancer (CACNA2D2/alpha2delta-2, SEMA3B [formerly SEMA(V), BLU, and HYAL1] or small cell lung cancer (SEMA3B, BLU, and HYAL1) cell lines. We found six of the genes to have two or more amino acid sequence-altering mutations including BLU, NPRL2/Gene21, FUS1, HYAL1, FUS2, and SEMA3B. However, none of the 19 genes tested for mutation showed a frequent (>10%) mutation rate in lung cancer samples. This led us to exclude several of the genes in the region as classical tumor suppressors for sporadic lung cancer. On the other hand, the putative lung cancer TSG in this location may either be inactivated by tumor-acquired promoter hypermethylation or belong to the novel class of haploinsufficient genes that predispose to cancer in a hemizygous (+/-) state but do not show a second mutation in the remaining wild-type allele in the tumor. We discuss the data in the context of novel and classic cancer gene models as applied to lung carcinogenesis. Further functional testing of the critical genes by gene transfer and gene disruption strategies should permit the identification of the putative lung cancer TSG(s), LUCA, Analysis of the approximately 630-kb sequence also provides an opportunity to probe and understand the genomic structure, evolution, and functional organization of this relatively gene-rich region.

Carcinoma, Non-Small-Cell Lung↗

Development of a pooled probe method for locating small gene families in a physical map of soybean using stress related paralogues and a BAC minimum tile path.

BACKGROUND: Genome analysis of soybean (Glycine max L.) has been complicated by its paleo-autopolyploid nature and conserved homeologous regions. Landmarks of expressed sequence tags (ESTs) located within a minimum tile path (MTP) of contiguous (contig) bacterial artificial chromosome (BAC) clones or radiation hybrid set can identify stress and defense related gene rich regions in the genome. A physical map of about 2,800 contigs and MTPs of 8,064 BAC clones encompass the soybean genome. That genome is being sequenced by whole genome shotgun methods so that reliable estimates of gene family size and gene locations will provide a useful tool for finishing. The aims here were to develop methods to anchor plant defense- and stress-related gene paralogues on the MTP derived from the soybean physical map, to identify gene rich regions and to correlate those with QTL for disease resistance. RESULTS: The probes included 143 ESTs from a root library selected by subtractive hybridization from a multiply disease resistant soybean cultivar 'Forrest' 14 days after inoculation with Fusarium solani f. sp. glycines (F. virguliforme). Another 166 probes were chosen from a root EST library (Gm-r1021) prepared from a non-inoculated soybean cultivar 'Williams 82' based on their homology to the known defense and stress related genes. Twelve and thirteen pooled EST probes were hybridized to high-density colony arrays of MTP BAC clones from the cv. 'Forrest' genome. The EST pools located 613 paralogues for 201 of the 309 probes used (range 1-13 per functional probe). One hundred BAC clones contained more than one kind of paralogue. Many more BACs (246) contained a single paralogue of one of the 201 probes detectable gene families. ESTs were anchored on soybean linkage groups A1, B1, C2, E, D1a+Q, G, I, M, H, and O. CONCLUSION: Estimates of gene family sizes were more similar to those made by Southern hybridization than by bioinformatics inferences from EST collections. When compared to Arabidopsis thaliana there were more 2 and 4 member paralogue families reflecting the diploidized-tetraploid nature of the soybean genome. However there were fewer families with 5 or more genes and the same number of single genes. Therefore the method can identify evolutionary patterns such as massively extensive selective gene loss or rapid divergence to regenerate the unique genes in some families.

Journal Article↗