Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Contig Mapping”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

A novel gene, MALT1 at 18q21, is involved in t(11;18) (q21;q21) found in low-grade B-cell lymphoma of mucosa-associated lymphoid tissue.

The t(11;18) (q21;q21) translocation is a characteristic chromosomal aberration in low-grade B-cell lymphoma of mucosa-associated lymphoid tissue (MALT) type. We previously identified a YAC clone y789F3, which includes the breakpoint at 18q21 in a MALT lymphoma patient. BAC and PAC contigs were constructed on the YAC, and BAC 193f9 was found to encompass the breakpoint region. In the present study, we further narrowed down the breakpoint region at 18q21 in five MALT lymphoma patients by means of FISH and Southern blot analyses using the plasmid contig constructed from BAC 193f9. The breakpoints at 18q21 in three of the five MALT lymphoma patients were found to be clustered approximately within the 20 kb region. By using exon amplification and cDNA library screening, we identified a novel cDNA spanning the breakpoint region that exhibited aberrant mRNA signals in four of the five MALT lymphoma patients. The nucleotide sequence predicted an 813 amino acid protein that shows significant sequence similarity to the CD22beta and laminin 5 alpha3b subunit. We refer to the gene encoding this transcript as MALT1 (Mucosa-Associated Lymphoid Tissue lymphoma translocation gene 1). The alteration of MALT1 by translocation strongly suggests that this gene plays an important role in the pathogenesis of MALT lymphoma.

Amino Acid Sequence↗

Identification of a YAC from 16q24 carrying a senescence gene for breast cancer cells.

We have identified a 360 kb YAC that carries a cell senescence gene, SEN16. In our earlier studies, we localized SEN16 within a genetic interval of 3 - 7 cM at 16q24.3. Six overlapping YACs spanning the chromosomal region of senescence activity, were assembled in a contig. Candidate YACs, identified by the markers located in the vicinity of SEN16, were retrofitted to introduce a neo selectable marker. Retrofitted YACs were first transferred into mouse A9 cells to generate A9/YAC hybrids. YAC DNA present in A9/YAC hybrids was further transferred by microcell fusion into immortal cell lines derived from human and rat mammary tumors. YAC d792t2 restored senescence in both human and rat mammary tumor cell lines, while an unrelated YAC from chromosome 6q had no senescence activity.

Aging↗

Aflatoxin B1 aldehyde reductase (AFAR) genes cluster at 1p35-1p36.1 in a region frequently altered in human tumour cells.

Alterations of the distal portion of the short arm of chromosome 1 (1p) are among the earliest abnormalities of human colorectal tumours. Recently, we have cloned the Aflatoxin B1 aldehyde reductase (AFAR) gene from a smallest region of overlapping deletion that is frequently (48%) hemizygously deleted in sporadic colorectal cancer. AFAR is expressed in a broad range of tissues. Its closely related rat protein is the major factor conferring resistance of rats towards aflatoxin B1-induced liver carcinogenesis. Here, we have identified cDNAs covering two additional human AFAR-related genes localized in close proximity to the previously described AFAR at 1p35-36. We have analysed their structure and tissue-related expression. One of them, AFAR3, carries a Selenocysteine-Insertion Element (SECIS)-like structure that during translation may recode an in-frame TGA-stop codon to a selenocysteine. Two additional AFAR-pseudogenes are localized at Xq25 and 1p12, respectively. AFAR exon sequences share an identity of DNA and amino acids of more than 78%. Also large blocks of intronic sequences can be up to 98.6% identical. Knowledge of the AFAR genes and their structure will be essential in genetic and functional studies, where discrimination of the genes and proteins is a prerequisite for evaluating their individual functions.

Aldehyde Reductase↗

A BAC contig of approximately 400 kb contains the classical class I major histocompatibility complex (MHC) genes of cattle.

A cattle BAC library derived from an MHC homozygous animal was screened for MHC class I genes. This revealed at least nine class I-related genes in a contig spanning approximately 400 kb, and several additional genes on other clones. The three classical class I genes expressed on this haplotype (A14) were shown to be distributed over a region at most 212 kb apart.

Animals↗

Comparison of EST libraries from seven beetle species: towards a framework for phylogenomics of the Coleoptera.

Relatively little is known about Coleoptera genes and genomes and how these compare in different taxa. We describe here the construction, DNA sequencing and sequence comparisons of cDNA libraries from seven beetle species. A total of 6717 bacterial colonies were screened for cDNA insert containing plasmids and 2784 size selected clones were 5'- and 3'-end sequenced to produce 1620 assembled sequences. Similarity comparisons with existing protein sequence databases revealed that 65.1% had matches (E < 10(-4)) in other organisms, with greater numbers of matches in Drosophila melanogaster than Caenorhabditis elegans and Saccharomyces cerevisiae databases. tBlastX comparisons also revealed numerous similarity hits (E < 10(-20)) in intra- and interlibrary comparisons. These results show the potential of small cDNA libraries for discovery and comparative analysis of genes useful for phylogenomic and functional studies.

Animals↗

The Arabidopsis PBS1 resistance gene encodes a member of a novel protein kinase subfamily.

Specific recognition of Pseudomonas syringae strains that express the avirulence gene avrPphB requires two genes in Arabidopsis, RPS5 and PBS1. Previous work has shown that RPS5 encodes a member of the nucleotide binding site-leucine rich repeat class of plant disease resistance genes. Here we report that PBS1 encodes a putative serine-threonine kinase. Southern blot analysis revealed that the pbs1-1 allele contained a deletion of the 3' end of the PBS1 open reading frame. DNA sequence analysis of the pbs1-2 allele showed it to be a missense mutation that caused a glycine to arginine substitution in the activation segment of PBS1, a region known to regulate substrate binding and catalytic activity in many protein kinases. The identity of PBS1 was confirmed using both transient transformation and stable transformation of mutant pbs1 plants. Comparison of the predicted PBS1 amino acid sequence with other plant protein kinases revealed that PBS1 belongs to a distinct subfamily of protein kinases that contains no other members of known function. The Pto kinase of tomato, which is required for specific resistance to P. syringae strains expressing avrPto, did not fall in the same subfamily as PBS1 and is only 42% identical in the kinase domain. These data suggest that PBS1 and Pto may fulfil different functions in the recognition of pathogen avirulence proteins. We discuss several possible models for the roles of PBS1 and RPS5 in AvrPphB recognition.

Alleles↗

Correlated clustering and virtual display of gene expression patterns in the wheat life cycle by large-scale statistical analyses of expressed sequence tags.

Compared to rice, wheat exhibits characteristic growth habits and contains complex genome constituents. To assess global changes in gene expression patterns in the wheat life cycle, we conducted large-scale analysis of expressed sequence tags (ESTs) in common wheat. Ten wheat tissues were used to construct cDNA libraries: crown and root from 14-day-old seedlings; spikelet from early and late flowering stages; spike at the booting stage, heading date and flowering date; pistil at the heading date; and seeds at 10 and 30 days post-anthesis. Several thousand colonies were randomly selected from each of these 10 cDNA libraries and sequenced from both 5' and 3' ends. Consequently, a total of 116 232 sequences were accumulated and classified into 25 971 contigs based on sequence homology. By computing abundantly expressed ESTs, correlated expression patterns of genes across the tissues were identified. Furthermore, relationships of gene expression profiles among the 10 wheat tissues were inferred from global gene expression patterns. Genes with similar functions were grouped with one another by clustering gene expression profiles. This technique might enable estimation of the functions of anonymous genes. Multidimensional analysis of EST data that is analogous to the microarray experiments may offer new approaches to functional genomics of plants.

Contig Mapping↗

Construction of a bacterial artificial chromosome (BAC) contig across the minimally deleted region in 13q14.3 in B-cell chronic lymphocytic leukemia.

Loss of heterozygosity (LOH) analysis in B-cell chronic lymphocytic leukemia (BCLL) has indicated that a frequent genetic event is loss of alleles from an approximately 500 kb region in 13q14.3, distal to the retinoblastoma gene. We have used DNA markers from this region to isolate and characterize a series of bacterial artificial chromosomes (BACs) which span the region between markers D13S319 and D13S25, which represents the common region of LOH. This entire region appears to be contained within only two minimally overlapping BACs, representing a maximum distance of approximately 350 kb. This BAC contig has been used to position known STS, EST and gene markers within the region. We have also used a modified differential display/RNA fingerprinting procedure designed to isolated transcribed sequences from YACs to isolate two transcribed units from the region which have also been positioned within the contig. The construction of a BAC contig with minimal redundancy provides the ideal resources from which to begin to identify candidate genes related to BCLL.

Blotting, Southern↗

The causal element for the lactase persistence/non-persistence polymorphism is located in a 1 Mb region of linkage disequilibrium in Europeans.

Expression of lactase in the intestine persists into adult life in some people and not others, and this is due to a cis-acting regulatory polymorphism. Previous data indicated that a mutation leading to lactase persistence had occurred on the background of a 60 kb 11-site LCT haplotype known as A (Hollox et al. 2001). Recent studies reported a 100% correlation of lactase persistence with the presence of the T allele at a CT SNP at -14 kb from LCT, in individuals of Finnish origin, suggesting that this SNP may be causal of the lactase persistence polymorphism, and also reported a very tight association with a second SNP (GA -22 kb) (Enattah et al. 2002). Here we report the existence of a one megabase stretch of linkage disequilibrium in the region of LCT and show that the -14 kb T allele and the -22 kb A allele both occur on the background of a very extended A haplotype. In a series of Finnish individuals we found a strong correlation (40/41 people) with lactose digestion and the presence of the T allele. The T allele was present in all 36 lactase persistent individuals from the UK (phenotyped by enzyme assay) studied, 31/36 of whom were of Northern European ancestry, but not in 11 non-persistent individuals who were mainly of non-UK ancestry. However, the CT heterozygotes did not show intermediate lactase enzyme activity, unlike those previously phenotyped by determining allelic transcript expression. Furthermore the one lactase persistent homozygote identified by having equally high expression of A and B haplotype transcripts, was heterozygous for CT at the -14 kb site. SNP analysis across the 1 megabase region in this person showed no evidence of recombination on either chromosome between the -14 kb SNP and LCT. The combined data shows that although the -14 kb CT SNP is an excellent candidate for the cause of the lactase persistence polymorphism, linkage disequilibrium extends far beyond the region searched so far. In addition, the CT SNP does not, on its own, explain all the variation in expression of LCT, suggesting the possibility of genetic heterogeneity.

Alleles↗

Long-range patterns of diversity and linkage disequilibrium surrounding the maize Y1 gene are indicative of an asymmetric selective sweep.

Both yellow and white corn occurs among ancestral open pollinated varieties. More recently, breeders have selected yellow endosperm variants of maize over ancestral white phenotypes for their increased nutritional value resulting from the up-regulation of the Y1 phytoene synthase gene product in endosperm tissue. As a result, diversity within yellow maize lines at the Y1 gene is dramatically decreased as compared to white corn. We analyzed patterns of sequence diversity and linkage disequilibrium in nine low copy regions located at varying distances from the Y1 gene, including a homolog of the barley Mlo gene. Patterns consistent with a selective sweep, such as significant associations of informative single-nucleotide polymorphisms with endosperm color phenotype, linkage disequilibrium, and significantly reduced diversity within the yellow endosperm haplotypes, were observed up to 600 kb downstream of Y1, whereas the upstream region showed a more rapid recovery. The starch branching enzyme 1 (sbe1) gene is the first region downstream of Y1 that does not have a highly conserved haplotype in the yellow endosperm germplasm.

Chromosomes, Artificial, Bacterial↗

Whole-genome shotgun assembly and comparison of human genome assemblies.

We report a whole-genome shotgun assembly (called WGSA) of the human genome generated at Celera in 2001. The Celera-generated shotgun data set consisted of 27 million sequencing reads organized in pairs by virtue of end-sequencing 2-kbp, 10-kbp, and 50-kbp inserts from shotgun clone libraries. The quality-trimmed reads covered the genome 5.3 times, and the inserts from which pairs of reads were obtained covered the genome 39 times. With the nearly complete human DNA sequence [National Center for Biotechnology Information (NCBI) Build 34] now available, it is possible to directly assess the quality, accuracy, and completeness of WGSA and of the first reconstructions of the human genome reported in two landmark papers in February 2001 [Venter, J. C., Adams, M. D., Myers, E. W., Li, P. W., Mural, R. J., Sutton, G. G., Smith, H. O., Yandell, M., Evans, C. A., Holt, R. A., et al. (2001) Science 291, 1304-1351; International Human Genome Sequencing Consortium (2001) Nature 409, 860-921]. The analysis of WGSA shows 97% order and orientation agreement with NCBI Build 34, where most of the 3% of sequence out of order is due to scaffold placement problems as opposed to assembly errors within the scaffolds themselves. In addition, WGSA fills some of the remaining gaps in NCBI Build 34. The early genome sequences all covered about the same amount of the genome, but they did so in different ways. The Celera results provide more order and orientation, and the consortium sequence provides better coverage of exact and nearly exact repeats.

Computational Biology↗

Genomic data support the hominoid slowdown and an Early Oligocene estimate for the hominoid-cercopithecoid divergence.

Several lines of indirect evidence suggest that hominoids (apes and humans) and cercopithecoids (Old World monkeys) diverged around 23-25 Mya. Importantly, although this range of dates has been used as both an initial assumption and as a confirmation of results in many molecular-clock analyses, it has not been critically assessed on its own merits. In this article we test the robusticity of the 23- to 25-Mya estimate with approximately 150,000 base pairs of orthologous DNA sequence data from two cercopithecoids and two hominoids by using quartet analysis. This method is an improvement over other estimates of the hominoid-cercopithecoid divergence because it incorporates two calibration points, one each within cercopithecoids and hominoids, and tests for a statistically appropriate model of molecular evolution. Most comparisons reject rate constancy in favor of a model incorporating two rates of evolution, supporting the "hominoid slowdown" hypothesis. By using this model of molecular evolution, the hominoid-cercopithecoid divergence is estimated to range from 29.2 to 34.5 Mya, significantly older than most previous analyses. Hominoid-cercopithecoid divergence dates of 23-25 Mya fall outside of the confidence intervals estimated, suggesting that as much as one-third of ape evolution has not been paleontologically sampled. Identifying stem cercopithecoids or hominoids from this period will be difficult because derived features that define crown catarrhines need not be present in early members of these lineages. More sites that sample primate habitats from the Oligocene of Africa are needed to better understand early ape and Old World monkey evolution.

Animals↗

Discovery of five conserved beta -defensin gene clusters using a computational search strategy.

The innate immune system includes antimicrobial peptides that protect multicellular organisms from a diverse spectrum of microorganisms. beta-Defensins comprise one important family of mammalian antimicrobial peptides. The annotation of the human genome fails to reveal the expected diversity, and a recent query of the draft sequence with the blast search engine found only one new beta-defensin gene (DEFB3). To define better the beta-defensin gene family, we adopted a genomics approach that uses hmmer, a computational search tool based on hidden Markov models, in combination with blast. This strategy identified 28 new human and 43 new mouse beta-defensin genes in five syntenic chromosomal regions. Within each syntenic cluster, the gene sequences and organization were similar, suggesting each cluster pair arose from a common ancestor and was retained because of conserved functions. Preliminary analysis indicates that at least 26 of the predicted genes are transcribed. These results demonstrate the value of a genomewide search strategy to identify genes with conserved structural motifs. Discovery of these genes represents a new starting point for exploring the role of beta-defensins in innate immunity.

Amino Acid Motifs↗

Quality assessment of maize assembled genomic islands (MAGIs) and large-scale experimental verification of predicted genes.

Recent sequencing efforts have targeted the gene-rich regions of the maize (Zea mays L.) genome. We report the release of an improved assembly of maize assembled genomic islands (MAGIs). The 114,173 resulting contigs have been subjected to computational and physical quality assessments. Comparisons to the sequences of maize bacterial artificial chromosomes suggest that at least 97% (160 of 165) of MAGIs are correctly assembled. Because the rates at which junction-testing PCR primers for genomic survey sequences (90-92%) amplify genomic DNA are not significantly different from those of control primers ( approximately 91%), we conclude that a very high percentage of genic MAGIs accurately reflect the structure of the maize genome. EST alignments, ab initio gene prediction, and sequence similarity searches of the MAGIs are available at the Iowa State University MAGI web site. This assembly contains 46,688 ab initio predicted genes. The expression of almost half (628 of 1,369) of a sample of the predicted genes that lack expression evidence was validated by RT-PCR. Our analyses suggest that the maize genome contains between approximately 33,000 and approximately 54,000 expressed genes. Approximately 5% (32 of 628) of the maize transcripts discovered do not have detectable paralogs among maize ESTs or detectable homologs from other species in the GenBank NR nucleotide/protein database. Analyses therefore suggest that this assembly of the maize genome contains approximately 350 previously uncharacterized expressed genes. We hypothesize that these "orphans" evolved quickly during maize evolution and/or domestication.

Chromosomes, Artificial, Bacterial↗

Combining two genomes in one cell: stable cloning of the Synechocystis PCC6803 genome in the Bacillus subtilis 168 genome.

Cloning the whole 3.5-megabase (Mb) genome of the photosynthetic bacterium Synechocystis PCC6803 into the 4.2-Mb genome of the mesophilic bacterium Bacillus subtilis 168 resulted in a 7.7-Mb composite genome. We succeeded in such unprecedented large-size cloning by progressively assembling and editing contiguous DNA regions that cover the entire Synechocystis genome. The strain containing the two sets of genome grew only in the B. subtilis culture medium where all of the cloning procedures were carried out. The high structural stability of the cloned Synechocystis genome was closely associated with the symmetry of the bacterial genome structure of the DNA replication origin (oriC) and its termination (terC) and the exclusivity of Synechocystis ribosomal RNA operon genes (rrnA and rrnB). Given the significant diversity in genome structure observed upon horizontal DNA transfer in nature, our stable laboratory-generated composite genome raised fundamental questions concerning two complete genomes in one cell. Our megasize DNA cloning method, designated megacloning, may be generally applicable to other genomes or genome loci of free-living organisms.

Bacillus subtilis↗

A Sanger/pyrosequencing hybrid approach for the generation of high-quality draft assemblies of marine microbial genomes.

Since its introduction a decade ago, whole-genome shotgun sequencing (WGS) has been the main approach for producing cost-effective and high-quality genome sequence data. Until now, the Sanger sequencing technology that has served as a platform for WGS has not been truly challenged by emerging technologies. The recent introduction of the pyrosequencing-based 454 sequencing platform (454 Life Sciences, Branford, CT) offers a very promising sequencing technology alternative for incorporation in WGS. In this study, we evaluated the utility and cost-effectiveness of a hybrid sequencing approach using 3730xl Sanger data and 454 data to generate higher-quality lower-cost assemblies of microbial genomes compared to current Sanger sequencing strategies alone.

Biotechnology↗

Comparative genomic analysis in the region of a major Plasmodium-refractoriness locus of Anopheles gambiae.

We have sequenced six overlapping clones from a library of bacterial artificial chromosome (BAC) clones derived from a laboratory strain of the mosquito, Anopheles gambiae, the major vector of human malaria in Africa. The resulting uninterrupted 528-kb sequence is from the 8C region of the mosquito 2R chromosome, at or very near the major refractoriness locus associated with melanotic encapsulation of parasites. This sequence represents the first extensive view of the mosquito genome structure encompassing 48 genes. Genomic comparison reveals that the majority of the orthologues are found in six microsyntenic clusters in Drosophila melanogaster. A BAC clone that is wholly contained within this region demonstrates the existence of a remarkable degree of local polymorphism in this species, which may prove important for its population structure and vectorial capacity.

Amino Acid Sequence↗

Complex mtDNA constitutes an approximate 620-kb insertion on Arabidopsis thaliana chromosome 2: implication of potential sequencing errors caused by large-unit repeats.

Previously conducted sequence analysis of Arabidopsis thaliana (ecotype Columbia-0) reported an insertion of 270-kb mtDNA into the pericentric region on the short arm of chromosome 2. DNA fiber-based fluorescence in situ hybridization analyses reveal that the mtDNA insert is 618 +/- 42 kb, approximately 2.3 times greater than that determined by contig assembly and sequencing analysis. Portions of the mitochondrial genome previously believed to be absent were identified within the insert. Sections of the mtDNA are repeated throughout the insert. The cytological data illustrate that DNA contig assembly by using bacterial artificial chromosomes tends to produce a minimal clone path by skipping over duplicated regions, thereby resulting in sequencing errors. We demonstrate that fiber-fluorescence in situ hybridization is a powerful technique to analyze large repetitive regions in the higher eukaryotic genomes and is a valuable complement to ongoing large genome sequencing projects.

Arabidopsis↗