Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Contig Mapping”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

Barnacle: an assembly algorithm for clone-based sequences of whole genomes.

We propose an assembly algorithm Barnacle for sequences generated by the clone-based approach. We illustrate our approach by assembling the human genome. Our novel method abandons the original physical-mapping-first framework. As we show, Barnacle more effectively resolves conflicts due to repeated sequences which is the main difficulty of the sequence assembly problem. In addition, we are able to detect inconsistencies in the underlying data. We present and compare our results on the December 2001 freeze of the public working draft of the human genome with NCBI's assembly (Build 28). The assembly of December 2001 freeze of the public working draft generated by Barnacle and the source code of Barnacle are available at (http://www.cs.rutgers.edu/~vchoi).

Algorithms↗

Gene for the human transmembrane-type protein tyrosine phosphatase H (PTPRH): genomic structure, fine-mapping and its exclusion as a candidate for Peutz-Jeghers syndrome.

Mutations in the serine/threonine kinase STK11 lead to Peutz-Jeghers syndrome (PJS) in a subset of affected individuals. Significant evidence for linkage to a second potential PJS disease locus on 19q13.4 has previously been described in one PJS family (PJS07). In the current study, we investigated this second locus for PJS gene candidates. We mapped the main candidate gene in this region, the gene for the transmembrane-type protein tyrosine phosphatase H (PTPRH), within 15 kb telomeric to the marker D19S880. We determined its genomic structure, and performed mutation analysis of all exons and the exon-intron junctions of the PTPRH gene in the PJS07 family. No disease causing mutation was identified in PTPRH in affected individuals, suggesting the existence of an as yet not identified gene on 19q13.4 as a second PJS gene.

Chromosomes, Human, Pair 19↗

Construction of physical maps from oligonucleotide fingerprints data.

A new algorithm for the construction of physical maps from hybridization fingerprints of short oligonucleotide probes has been developed. Extensive simulations in high-noise scenarios show that the algorithm produces an essentially completely correct map in over 95% of trials. Tests for the influence of specific experimental parameters demonstrate that the algorithm is robust to both false positive and false negative experimental errors. The algorithm was also tested in simulations using real DNA sequences of C. elegans, E. coli, S. cerevisiae, and H. sapiens. To overcome the non-randomness of probe frequencies in these sequences, probes were preselected based on sequence statistics and a screening process of the hybridization data was developed. With these modifications, the algorithm produced very encouraging results.

Algorithms↗

The SPCH1 region on human 7q31: genomic characterization of the critical interval and localization of translocations associated with speech and language disorder.

The KE family is a large three-generation pedigree in which half the members are affected with a severe speech and language disorder that is transmitted as an autosomal dominant monogenic trait. In previously published work, we localized the gene responsible (SPCH1) to a 5.6-cM region of 7q31 between D7S2459 and D7S643. In the present study, we have employed bioinformatic analyses to assemble a detailed BAC-/PAC-based sequence map of this interval, containing 152 sequence tagged sites (STSs), 20 known genes, and >7.75 Mb of completed genomic sequence. We screened the affected chromosome 7 from the KE family with 120 of these STSs (average spacing <100 kb), but we did not detect any evidence of a microdeletion. Novel polymorphic markers were generated from the sequence and were used to further localize critical recombination breakpoints in the KE family. This allowed refinement of the SPCH1 interval to a region between new markers 013A and 330B, containing approximately 6.1 Mb of completed sequence. In addition, we have studied two unrelated patients with a similar speech and language disorder, who have de novo translocations involving 7q31. Fluorescence in situ hybridization analyses with BACs/PACs from the sequence map localized the t(5;7)(q22;q31.2) breakpoint in the first patient (CS) to a single clone within the newly refined SPCH1 interval. This clone contains the CAGH44 gene, which encodes a brain-expressed protein containing a large polyglutamine stretch. However, we found that the t(2;7)(p23;q31.3) breakpoint in the second patient (BRD) resides within a BAC clone mapping >3.7 Mb distal to this, outside the current SPCH1 critical interval. Finally, we investigated the CAGH44 gene in affected individuals of the KE family, but we found no mutations in the currently known coding sequence. These studies represent further steps toward the isolation of the first gene to be implicated in the development of speech and language.

Base Sequence↗

Mapping and initial analysis of human subtelomeric sequence assemblies.

Physical mapping data were combined with public draft and finished sequences to derive subtelomeric sequence assemblies for each of the 41 genetically distinct human telomere regions. Sequence gaps that remain on the reference telomeres are generally small,well-defined,and for the most part,restricted to regions directly adjacent to the terminal (TTAGGG)n tract. Of the 20.66 Mb of subtelomeric DNA analyzed, 3.01 Mb are subtelomeric repeat sequences (Srpt),and an additional 2.11 Mb are segmental duplications. The subtelomeric sequence assemblies are enriched >25-fold in short,internal (TTAGGG)n-like sequences relative to the rest of the genome; a total of 114 (TTAGGG)n-like islands were found,55 within Srpt regions,35 within one-copy regions,11 at one-copy/Srpt or Srpt/segmental duplication boundaries,and 13 at the telomeric ends of assemblies. Transcripts were annotated in each assembly,noting their mapping coordinates relative to their respective telomere and whether they originate in duplicated DNA or single-copy DNA. A total of 697 transcripts were found in 15.53 Mb of one-copy DNA,76 transcripts in 2.11 Mb of segmentally duplicated DNA,and 168 transcripts in 3.01 Mb of Srpt sequence. This overall transcript density is similar (within approximately 10%) to that found genome-wide. Zinc finger-containing genes and olfactory receptor genes are duplicated within and between multiple telomere regions.

Base Composition↗

Mapping and characterization of the mouse and human SS18 genes, two human SS18-like genes and a mouse Ss18 pseudogene.

We have previously isolated and characterized a mouse cDNA orthologous to the human synovial sarcoma associated SS18 (formerly named SSXT and SYT) cDNA. Here, we report the characterization of the genomic structure of the mouse Ss18 gene. Through in silico methods with sequence information contained in the public databases, we did the same for the human SS18 gene and two human SS18 homologous genes, SS18L1 and SS18L2. In addition, we identified a mouse Ss18 processed pseudogene and mapped it to chromosome 1, band A2-3. The mouse Ss18 gene, which is subject to extensive alternative splicing, is made up of 11 exons, spread out over approximately 45 kb of genomic sequence. The human SS18 gene is also composed of 11 exons with similar intron-exon boundaries, spreading out over about 70 kb of genomic sequence. One alternatively spliced exon, which is not included in the published SS18 cDNA, corresponds to a stretch of sequence which we previously identified in the mouse Ss18 cDNA. The human SS18L1 gene, which is also made up of 11 exons with similar intron-exon boundaries, was mapped to chromosome 20 band q13.3. The smaller SS18L2 gene, which is composed of three exons with similar boundaries as the first three exons of the other three genes, was mapped to chromosome 3 band p21. Through sequence and mutation analyses this gene could be excluded as a candidate gene for 3p21-associated renal cell cancer. In addition, we created a detailed BAC map around the human SS18 gene, placing it unequivocally between the CA-repeat marker AFMc014wf9 and the dihydrofolate reductase pseudogene DHFRP1. The next gene in this map, located distal to SS18, was found to be the TBP associated factor TAFII-105 (TAF2C2). Further analogies between the mouse Ss18 gene, the human SS18 gene and its two homologous genes were found in the putative promoter fragments. All four promoters resemble the promoters of housekeeping genes in that they are TATA-less and embedded in canonical CpG islands, thus explaining the high and widespread expression of the SS18 genes.

Alternative Splicing↗

Cloning, mapping, and expression of ial, a novel Drosophila member of the Ipl1/aurora mitotic control kinase family.

The members of the Ipl1-aurora like kinase (IARK) subfamily are conserved serine/threonine kinases that play a key role in the control of chromosome segregation, centrosome separation, and cytokinesis from yeast to mammals. We report on the isolation of a new Drosophila member of the family, designated Ipl1-aurora-like kinase (ial) Phylogenetic analysis of kinase domains established that ial is more divergent from known mammalian IARKs than is aurora. Mapping based on examination of chromosomal aberrations, together with mapping within contigs identified by the Drosophila Genome Project, placed the gene at 32B on the left arm of the second chromosome. Discrete single-gene mutations in this region, including all known relevant P-element disruptions, were examined and proven not to be mutations in ial. Characterization of spatial and temporal expression of ial and its gene product showed that it manifests itself in patterns which can be consistent with a role in cell cycle control.

Amino Acid Sequence↗

An integrated physical and genetic map of the rice genome.

Rice was chosen as a model organism for genome sequencing because of its economic importance, small genome size, and syntenic relationship with other cereal species. We have constructed a bacterial artificial chromosome fingerprint-based physical map of the rice genome to facilitate the whole-genome sequencing of rice. Most of the rice genome ( approximately 90.6%) was anchored genetically by overgo hybridization, DNA gel blot hybridization, and in silico anchoring. Genome sequencing data also were integrated into the rice physical map. Comparison of the genetic and physical maps reveals that recombination is suppressed severely in centromeric regions as well as on the short arms of chromosomes 4 and 10. This integrated high-resolution physical map of the rice genome will greatly facilitate whole-genome sequencing by helping to identify a minimum tiling path of clones to sequence. Furthermore, the physical map will aid map-based cloning of agronomically important genes and will provide an important tool for the comparative analysis of grass genomes.

Chromosomes, Artificial, Bacterial↗

Unequal VH gene rearrangement frequency within the large VH7183 gene family is not due to recombination signal sequence variation, and mapping of the genes shows a bias of rearrangement based on chromosomal location.

Much of the nonrandom usage of V, D, and J genes in the Ab repertoire is due to different frequencies with which gene segments undergo V(D)J rearrangement. The recombination signal sequences flanking each segment are seldom identical with consensus sequences, and this natural variation in recombination signal sequence (RSS) accounts for some differences in rearrangement frequencies in vivo. Here, we have sequenced the RSS of 19 individual V(H)7183 genes, revealing that the majority have one of two closely related RSS. One group has a consensus heptamer, and the other has a nonconsensus heptamer. In vitro recombination substrate studies show that the RSS with the nonconsensus heptamer, which include the frequently rearranging 81X, rearrange less well than the RSS with the consensus heptamer. Although 81X differs from the other 7183-I genes at three positions in the spacer, this does not significantly increase its recombination potency in vitro. The rearrangement frequency of all members of the family was determined in microMT mice, and there was no correlation between the in vitro recombination potential and V(H) gene rearrangement frequency in vivo. Furthermore, genes with identical RSS rearrange at different frequencies in vivo. This demonstrates that other factors can override differences in RSS potency in vivo. We have also determined the gene order of all V(H)7183 genes in a bacterial artificial chromosome contig and show that most of the frequently rearranging genes are in the 3' half of the region. This suggests that chromosomal location plays an important role in nonrandom rearrangement of the V(H)7183 genes.

Animals↗

Large-scale transcriptional activity in chromosomes 21 and 22.

The sequences of the human chromosomes 21 and 22 indicate that there are approximately 770 well-characterized and predicted genes. In this study, empirically derived maps identifying active areas of RNA transcription on these chromosomes have been constructed with the use of cytosolic polyadenylated RNA obtained from 11 human cell lines. Oligonucleotide arrays containing probes spaced on average every 35 base pairs along these chromosomes were used. When compared with the sequence annotations available for these chromosomes, it is noted that as much as an order of magnitude more of the genomic sequence is transcribed than accounted for by the predicted and characterized exons.

Cell Line↗

Genomic structure and paralogous regions of the inversion breakpoint occurring between human chromosome 3p12.3 and orangutan chromosome 2.

Intrachromosomal duplications play a significant role in human genome pathology and evolution. To better understand the molecular basis of evolutionary chromosome rearrangements, we performed molecular cytogenetic and sequence analyses of the breakpoint region that distinguishes human chromosome 3p12.3 and orangutan chromosome 2. FISH with region-specific BAC clones demonstrated that the breakpoint-flanking sequences are duplicated intrachromosomally on orangutan 2 and human 3q21 as well as at many pericentromeric and subtelomeric sites throughout the genomes. Breakage and rearrangement of the human 3p12.3-homologous region in the orangutan lineage were associated with a partial loss of duplicated sequences in the breakpoint region. Consistent with our FISH mapping results, computational analysis of the human chromosome 3 genomic sequence revealed three 3p12.3-paralogous sequence blocks on human chromosome 3q21 and smaller blocks on the short arm end 3p26-->p25. This is consistent with the view that sequences from an ancestral site at 3q21 were duplicated at 3p12.3 in a common ancestor of orangutan and humans. Our results show that evolutionary chromosome rearrangements are associated with microduplications and microdeletions, contributing to the DNA differences between closely related species.

Animals↗

Gene amplification in PNETs/medulloblastomas: mapping of a novel amplified gene within the MYCN amplicon.

OBJECTIVES: The pathological entity of primitive neuroectodermal tumour/medulloblastoma (PNET/MB) comprises a very heterogeneous group of neoplasms on a clinical as well as on a molecular level. We evaluated the importance of DNA amplification in medulloblastomas and other primitive neuroectodermal tumours (PNETs) of the CNS. METHOD: Restriction landmark genomic scanning (RLGS), a method that allows the detection of low level amplification, was used. RLGS provides direct access to DNA sequences circumventing positional cloning efforts. Furthermore, we analysed several samples by CGH. DESIGN: Twenty primary medulloblastomas, five supratentorial PNETs, and five medulloblastoma cell lines were studied. RESULTS: Although our analysis confirms that gene amplification is generally a rare event in childhood PNET/MB, we found a total of 17 DNA fragments that were amplified in seven different tumours. Cloning and sequencing of several of these fragments confirmed the previous finding of MYC amplification in the cell line D341 Med and identified novel DNA sequences amplified in PNET/MB. We describe for the first time amplification of the novel gene, NAG, in a subset of PNET/MB. Despite genomic amplification, NAG was not overexpressed in the tumours studied. We have determined that NAG maps less than 50 kb 5' of DDX1 and approximately 400 kb telomeric of MYCN on chromosome 2p24. CONCLUSION: We found a similar but slightly higher frequency of amplification than previously reported. We present several DNA fragments that may belong to the CpG islands of novel genes amplified in a small subset of PNET/MB. As an example we describe for the first time the amplification of NAG in the MYCN amplicon in PNET/MB.

Blotting, Northern↗

FISH analysis of terminal deletions in patients diagnosed with cri-du-chat syndrome.

Most patients with cri-du-chat syndrome have a de novo deletion of the short arm of chromosome 5 (5p). In order to perform extensive phenotype-genotype correlation studies, a relatively easy method for the precise determination of the extent of a patient's deletion is essential. Towards this purpose, a set of minimally overlapping YAC clones that span 5p was identified. A BAC that maps at or near the 5p telomere was also used. A total of 110 patients with previously determined de novo terminal deletions by standard cytogenetic approaches were reanalyzed using the YAC clones and fluorescent in situ hybridization (FISH). Of the 110 samples, 4 patients were determined to have interstitial deletions, 1 patient had an unbalanced translocation, and no deletion could be detected in 2 patients. The FISH results in the 7 patients affect the clinical prognosis for some of these patients. These results demonstrate the need for supplementing standard cytogenetics with FISH analysis when an abnormal karyotype is detected.

Child↗

Physical map and characterization of transcripts in the candidate interval for familial chondrocalcinosis at chromosome 5p15.1.

The gene for familial chondrocalcinosis (MIM 118600; gene symbol CCAL2) has been localized to a 0.8-cM interval on the short arm of chromosome 5, between the polymorphic microsatellite markers D5S416 and D5S2114. We have undertaken the physical and transcript mapping of this interval, as well as regions telomeric to the interval, in an attempt to define ultimately the gene for this disorder. The physical map is composed of YAC, BAC, PAC, and cosmid resources and spans a physical distance of approximately 0.3 Mb. Using cDNA selection, we have identified eight novel transcripts in and around the interval; two of the selected transcripts reside in the candidate interval. We have also more precisely placed several expressed sequence tags (ESTs) that were previously mapped by radiation hybrid analysis and were reported to reside in or near the candidate interval. Two of the ESTs analyzed overlap with the selected cDNAs that reside in the candidate interval. All of the selected cDNAs are expressed partial transcripts, as determined by Northern blot analysis, and using RT-PCR analysis, we have determined that the cDNAs that reside in the candidate interval are expressed in cartilage and synovium, tissues that are presumably relevant to the chondrocalcinosis phenotype.

Adult↗

Structure and mutation analysis of the gene encoding DNA fragmentation factor 40 (caspase-activated nuclease), a candidate neuroblastoma tumour suppressor gene.

We have characterised the DFFB gene, encoding the active subunit of the apoptotic nuclease DNA fragmentation factor (DFF40). DFFB maps to 1p36, near the imprinted putative tumour suppressor gene TP73. The DFFA gene (encoding the inhibitory DFF45 subunit) also maps to 1p36.2-36.3, and we show by FISH that DFFB lies distal to DFFA. We have also mapped a processed DFFB pseudogene to chromosome 9. DFFB itself has seven coding exons spanning 10 kb. Exhaustive mutation screening of 41 neuroblastomas and other tumours in which a 1p36 tumour suppressor gene is implicated showed no tumour-specific mutations. A coding region polymorphism was used to demonstrate uniformly biallelic expression in human fetal DFFB transcripts. Since the putative neuroblastoma tumour suppressor gene in distal 1p36 is predicted to be maternally expressed, the lack of imprinting and absence of somatic mutations in DFFB indicate that it is probably not the neuroblastoma tumour suppressor gene.

Apoptosis↗

Molecular cytogenetics and DNA sequence analysis of an apomixis-linked BAC in Paspalum simplex reveal a non pericentromere location and partial microcolinearity with rice.

Apomixis in plants is a form of clonal reproduction through seeds. A BAC clone linked to apomictic reproduction in Paspalum simplex was used to locate the apomixis locus on meiotic chromosome preparations. Fluorescent in situ hybridisation revealed the existence of a single locus embedded in a heterochromatin-poor region not adjacent to the centromere. We report here for the first time information regarding the sequencing of a large DNA clone from the apomixis locus. The presence of two genes whose rice homologs were mapped on the telomeric part of the long arm of rice chromosome 12 confirmed the strong synteny between the apomixis locus of P. simplex with the related area of the rice genome at the map level. Comparative analysis of this region with rice as representative of a sexual species revealed large-scale rearrangements due to transposable elements and small-scale rearrangements due to deletions and single point mutations. Both types of rearrangements induced the loss of coding capacity of large portions of the "apomictic" genes compared to their rice homologs. Our results are discussed in relation to the use of rice genome data for positional cloning of apomixis genes and to the possible role of rearranged supernumerary genes in the apomictic process of P. simplex.

Chromosomes, Artificial, Bacterial↗

Significant microsynteny with new evolutionary highlights is detected between Arabidopsis and legume model plants despite the lack of macrosynteny.

The increased amount of data produced by large genome sequencing projects allows scientists to carry out important syntenic studies to a great extent. Detailed genetic maps and entirely or partially sequenced genomes are compared, and macro- and microsyntenic relations can be determined for different species. In our study, the syntenic relationships between key legume plants and two model plants, Arabidopsis thaliana and Populus trichocarpa were investigated. The comparison of the map position of 172 gene-based Medicago sativa markers to the organization of homologous A. thaliana genes could not identify any sign of macrosynteny between the two genomes. A 276 kb long section of chromosome 5 of the model legume Medicago truncatula was used to investigate potential microsynteny with the other legume Lotus japonicus, as well as with Arabidopsis and Populus. Besides the overall correlation found between the legume plants, the comparison revealed several microsyntenic regions in the two more distant plants with significant resemblance. Despite the large phylogenetic distance, clear microsyntenic regions between Medicago and Arabidopsis or Populus were detected unraveling new intragenomic evolutionary relations in Arabidopsis.

Arabidopsis↗

Homozygous mutations in ARIX(PHOX2A) result in congenital fibrosis of the extraocular muscles type 2.

Isolated strabismus affects 1-5% of the general population. Most forms of strabismus are multifactorial in origin; although there is probably an inherited component, the genetics of these disorders remain unclear. The congenital fibrosis syndromes (CFS) represent a subset of monogenic isolated strabismic disorders that are characterized by restrictive ophthalmoplegia, and include congenital fibrosis of the extraocular muscles (CFEOM) and Duane syndrome (DURS). Neuropathologic studies indicate that these disorders may result from the maldevelopment of the oculomotor (nIII), trochlear (nIV) and abducens (nVI) cranial nerve nuclei. To date, five CFS loci have been mapped (FEOM1, FEOM2, FEOM3, DURS1 and DURS2), but no genes have been identified. Here, we report three mutations in ARIX (also known as PHOX2A) in four CFEOM2 pedigrees. ARIX encodes a homeodomain transcription factor protein previously shown to be required for nIII/nIV development in mouse and zebrafish. Two of the mutations are predicted to disrupt splicing, whereas the third alters an amino acid within the conserved brachyury-like domain. These findings confirm the hypothesis that CFEOM2 results from the abnormal development of nIII/nIV (ref. 7) and emphasize a critical role for ARIX in the development of these midbrain motor nuclei.

Amino Acid Sequence↗