Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Contig Mapping”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

Large-scale evaluation of imprinting status in the Prader-Willi syndrome region: an imprinted direct repeat cluster resembling small nucleolar RNA genes.

Loss of paternal gene expression at the imprinted domain on proximal human chromosome 15 causes Prader-Willi syndrome (PWS), a complex multiple-anomaly disorder involving variable mental retardation, hyperphasia leading to obesity and infantile hypotonia with failure to thrive. Although numerous paternally expressed transcripts have been identified that reside in the candidate region, the individual contributions to the development of PWS have not been firmly established. Recent studies of mouse models carrying a cytogenetic deletion suggest that paternal deficiency of the SNRPN-IPW interval is critical for perinatal lethality of potential relevance to PWS. Here we determined the allelic expression profiles of a total of 118 cDNA clones using monochromosomal hybrids retaining either a paternal or maternal human chromosome 15. Our results demonstrated a preponderance of unusual transcripts lacking protein-coding potential that were expressed exclusively from the paternal copy of the critical interval. This interval was also found to encompass a large direct repeat (DR) cluster displaying a potentially active chromatin conformation of paternal origin, as suggested by enhanced sensitivity to nuclease digestion. Database searches revealed an unexpected organization of tandemly repeated consensus elements, all of which possessed well-defined box C and D sequences characteristic of small nucleolar RNAs (snoRNAs). Southern blot analysis further demonstrated a considerable degree of phylogenetic conservation of the DR locus in the genomes of all mammalian species tested, but not in chicken, Xenopus and Drosophila. These findings imply a potential direct contribution of the DR locus, representing a cluster of multiple snoRNA genes, to certain phenotypic features of PWS.

Base Sequence↗

Syntenic organization of the mouse distal chromosome 7 imprinting cluster and the Beckwith-Wiedemann syndrome region in chromosome 11p15.5.

In human and mouse, most imprinted genes are arranged in chromosomal clusters. Their linked organization suggests co-ordinated mechanisms controlling imprinting and gene expression. The identification of local and regional elements responsible for the epigenetic control of imprinted gene expression will be important in understanding the molecular basis of diseases associated with imprinting such as Beckwith-Wiedemann syndrome. We have established a complete contig of clones along the murine imprinting cluster on distal chromosome 7 syntenic with the human imprinting region at 11p15.5 associated with Beckwith-Wiedemann syndrome. The cluster comprises approximately 1 Mb of DNA, contains at least eight imprinted genes and is demarcated by the two maternally expressed genes Tssc3 (Ipl) and H19 which are directly flanked by the non-imprinted genes Nap1l4 (Nap2) and Rpl23l (L23mrp), respectively. We also localized Kcnq1 (Kvlqt1) and Cd81 (Tapa-1) between Cdkn1c (p57(Kip2)) and Mash2. The mouse Kcnq1 gene is maternally expressed in most fetal but biallelically transcribed in most neonatal tissues, suggesting relaxation of imprinting during development. Our findings indicate conserved control mechanisms between mouse and human, but also reveal some structural and functional differences. Our study opens the way for a systematic analysis of the cluster by genetic manipulation in the mouse which will lead to animal models of Beckwith-Wiedemann syndrome and childhood tumours.

Amino Acid Sequence↗

The beta3A subunit gene (Ap3b1) of the AP-3 adaptor complex is altered in the mouse hypopigmentation mutant pearl, a model for Hermansky-Pudlak syndrome and night blindness.

Lysosomes, melanosomes and platelet-dense granules are abnormal in the mouse hypopigmentation mutant pearl. The beta3A subunit of the AP-3 adaptor complex, which likely regulates protein trafficking in the trans - Golgi network/endosomal compartments, was identified as a candidate for the pearl gene by a positional/candidate cloning approach. Mutations, including a large internal tandem duplication and a deletion, were identified in two respective pearl alleles and are predicted to abrogate function of the beta3A protein. Significantly lowered expression of altered beta3A transcripts occurred in kidney of both mutant alleles. The several distinct pearl phenotypes suggest novel functions for the AP-3 complex in mammals. These experiments also suggest mutations in AP-3 subunits as a basis for unique forms of human Hermansky-Pudlak syndrome and congenital night blindness, for which the pearl mouse is an appropriate animal model.

Adaptor Protein Complex beta Subunits↗

Chromosome 22-specific low copy repeats and the 22q11.2 deletion syndrome: genomic organization and deletion endpoint analysis.

The 22q11.2 deletion syndrome, which includes DiGeorge and velocardiofacial syndromes (DGS/VCFS), is the most common microdeletion syndrome. The majority of deleted patients share a common 3 Mb hemizygous deletion of 22q11.2. The remaining patients include those who have smaller deletions that are nested within the 3 Mb typically deleted region (TDR) and a few with rare deletions that have no overlap with the TDR. The identification of chromosome 22-specific duplicated sequences or low copy repeats (LCRs) near the end-points of the 3 Mb TDR has led to the hypothesis that they mediate deletions of 22q11.2. The entire 3 Mb TDR has been sequenced, permitting detailed investigation of the LCRs and their involvement in the 22q11.2 deletions. Sequence analysis has identified four LCRs within the 3 Mb TDR. Although the LCRs differ in content and organization of shared modules, those modules that are common between them share 97-98% sequence identity with one another. By fluorescence in situ hybridization (FISH) analysis, the end-points of four variant 22q11.2 deletions appear to localize to the LCRs. Pulsed-field gel electrophoresis and Southern hybridization have been used to identify rearranged junction fragments from three variant deletions. Analysis of junction fragments by PCR and sequencing of the PCR products implicate the LCRs directly in the formation of 22q11.2 deletions. The evolutionary origin of the duplications on chromosome 22 has been assessed by FISH analysis of non-human primates. Multiple signals in Old World monkeys suggest that the duplication events may have occurred at least 20-25 million years ago.

Animals↗

Duplicated copies of the bovine JH locus contribute to the Ig repertoire.

We report the cloning and analysis of a bovine JH locus comprising a DQ52 segment, six JH segments and sequence to a 5' H chain intronic enhancer. The contig was mapped to BTA 11 and evidence was found for rearrangement of the sixth JH segment at a low but detectable frequency. In contrast, the fourth segment present at a second copy of the bovine JH locus mapping to BTA 21 was found to rearrange at high-frequency, forming FR4 in the majority of bovine Ig H chains. The data thus show that bovine H chains can be generated from segments at two distinct genomic locations. Further investigation should establish if rearrangement takes place at each locus or if the participating segments are brought together from different chromosomal locations by less conventional processes (for example by gene conversion or trans-chromosomal rearrangement).

Amino Acid Sequence↗

Signal sequence conservation and mature peptide divergence within subgroups of the murine beta-defensin gene family.

Beta-defensins are two exon genes which encode broad spectrum antimicrobial cationic peptides. We have analyzed the largest murine cluster of these genes which localizes to chromosome 8. Using hidden Markov models, we identified six beta-defensin exon 2-like sequences and subsequently found full-length expressed transcripts for these novel genes. Expression was high in brain and reproductive tissues. Eleven beta-defensins could be grouped into two clear subgroups by virtue of their position and high signal sequence (exon 1 encoded) identity. In contrast, however, there was a very low level of sequence conservation in the exon 2 region encoding the mature antimicrobial peptide. Examination of the gene sequences of orthologs in other rodents also revealed an excess of nucleotide changes that altered amino acids in the mature peptide region. Evolutionary analysis revealed strong evidence that following gene duplication, exon 1 and surrounding noncoding DNA show little divergence within subgroups. The focus for rapid sequence divergence is localized in the DNA encoding the mature peptide and this is driven by accelerated positive selection. This mechanism of evolution is consistent with the role of this gene family as defense against bacterial pathogens and the sequence changes have implications for novel antibiotic design.

Animals↗

High frequency of DAZ1/DAZ2 gene deletions in patients with severe oligozoospermia.

Deletions of the DAZ gene family in distal Yq11 are always associated with deletions of the azoospermia factor c (AZFc) region, which we now estimate extends to 4.94 Mb. Because more Y gene families are located in this chromosomal region, and are expressed like the DAZ gene family only in the male germ line, the testicular pathology associated with complete AZFc deletions cannot predict the functional contribution of the DAZ gene family to human spermatogenesis. We therefore established a DAZ gene copy specific deletion analysis based on the DAZ-BAC sequences in GenBank. It includes the deletion analysis of eight DAZ-DNA PCR markers [six DAZ-single nucleotide varients (SNVs) and two DAZ-sequence tag sites (STS)] selected from the 5' to the 3'end of each DAZ gene and a deletion analysis of the gene copy specific EcoRV and TaqI restriction fragments identified in the internal repetitive DAZ gene regions (DYS1 locus). With these diagnostic tools, 63 DNA samples from men with idiopathic oligozoospermia and 107 DNA samples from men with proven fertility were analysed for the presence of the complete DAZ gene locus, encompassing the four DAZ gene copies. In five oligozoospermic patients, we found a DAZ-SNV/STS and DYS1/EcoRV and TaqI fragment deletion pattern indicative for deletion of the DAZ1 and DAZ2 gene copies; one of these deletions could be identified as a 'de-novo' deletion because it was absent in the DAZ locus of the patient's father. The same DAZ deletions were not found in any of the 107 fertile control samples. We therefore conclude that the deletion of the DAZ1/DAZ2 gene doublet in five out of our 63 oligozoospermic patients (8%) is responsible for the patients' reduced sperm numbers. It is most likely caused by intrachromosomal recombination events between two long repetitive sequence blocks (AZFc-Rep1) flanking the DAZ gene structures.

Chromosomes, Artificial, Bacterial↗

Interspecies conservation of gene order and intron-exon structure in a genomic locus of high gene density and complexity in Plasmodium.

A 13.6 kb contig of chromosome 5 of Plasmodium berghei, a rodent malaria parasite, has been sequenced and analysed for its coding potential. Assembly and comparison of this genomic locus with the orthologous locus on chromosome 10 of the human malaria Plasmodium falciparum revealed an unexpectedly high level of conservation of the gene organisation and complexity, only partially predicted by current gene-finder algorithms. Adjacent putative genes, transcribed from complementary strands, overlap in their untranslated regions, introns and exons, resulting in a tight clustering of both regulatory and coding sequences, which is unprecedented for genome organisation of PLASMODIUM: In total, six putative genes were identified, three of which are transcribed in gametocytes, the precursor cells of gametes. At least in the case of two multiple exon genes, alternative splicing and alternative transcription initiation sites contribute to a flexible use of the dense information content of this locus. The data of the small sample presented here indicate the value of a comparative approach for Plasmodium to elucidate structure, organisation and gene content of complex genomic loci and emphasise the need to integrate biological data of all Plasmodium species into the P.falciparum genome database and associated projects such as PlasmodB to further improve their annotation.

Alternative Splicing↗

The Dictyostelium discoideum family of Rho-related proteins.

Taking advantage of the ongoing Dictyostelium genome sequencing project, we have assembled >73 kb of genomic DNA in 15 contigs harbouring 15 genes and one pseudogene of Rho-related proteins. Comparison with EST sequences revealed that every gene is interrupted by at least one and up to four introns. For racC extensive alternative splicing was identified. Northern blot analysis showed that mRNAs for racA, racE, racG, racH and racI were present at all stages of development, whereas racJ and racL were expressed only at late stages. Amino acid sequences have been analysed in the context of Rho-related proteins of other organisms. Rac1a/1b/1c, RacF1/F2 and to a lesser extent RacB and the GTPase domain of RacA can be grouped in the Rac subfamily. None of the additional Dictyostelium Rho-related proteins belongs to any of the well-defined subfamilies, like Rac, Cdc42 or Rho. RacD and RacA are unique in that they lack the prenylation motif characteristic of Rho proteins. RacD possesses a 50 residue C-terminal extension and RacA a 400 residue C-terminal extension that contains a proline-rich region, two BTB domains and a novel C-terminal domain. We have also identified homologues for RacA in Drosophila and mammals, thus defining a new subfamily of Rho proteins, RhoBTB.

Alternative Splicing↗

Fast algorithms for large-scale genome alignment and comparison.

We describe a suffix-tree algorithm that can align the entire genome sequences of eukaryotic and prokaryotic organisms with minimal use of computer time and memory. The new system, MUMmer 2, runs three times faster while using one-third as much memory as the original MUMmer system. It has been used successfully to align the entire human and mouse genomes to each other, and to align numerous smaller eukaryotic and prokaryotic genomes. A new module permits the alignment of multiple DNA sequence fragments, which has proven valuable in the comparison of incomplete genome sequences. We also describe a method to align more distantly related genomes by detecting protein sequence homology. This extension to MUMmer aligns two genomes after translating the sequence in all six reading frames, extracts all matching protein sequences and then clusters together matches. This method has been applied to both incomplete and complete genome sequences in order to detect regions of conserved synteny, in which multiple proteins from one organism are found in the same order and orientation in another. The system code is being made freely available by the authors.

Algorithms↗

Cassette-like variation of restriction enzyme genes in Escherichia coli C and relatives.

A surprising result of comparative bacterial genomics has been the large amount of DNA found to be present in one strain but not in another of the same species. We examine in detail one location where gene content varies extensively, the restriction cluster in Escherichia coli. This region is designated the Immigration Control Region (ICR) for the density and variability of restriction functions found there. To better define the boundaries of this variable locus, we determined the sequence of the region from a restrictionless strain, E.coli C. Here we compare the 13.7 kb E.coli C sequence spanning the site of the ICR with corresponding sequences from five E.coli strains and Salmonella typhimurium LT2. To discuss this variation, we adopt the term 'framework' to refer to genes that are stable components of genomes within related lineages, while 'migratory' genes are transient inhabitants of the genome. Strikingly, seven different migratory DNA segments, encoding different sets of genes and gene fragments, alternatively occupy a single well-defined location in the seven strains examined. The flanking framework genes, yjiS and yjiA, display approximately normal patterns of conservation. The patterns observed are consistent with the action of a site-specific recombinase. Since no nearby gene codes for a likely recombinase of known families, such a recombinase must be of a new family or unlinked.

Base Sequence↗

An expressed sequence tag analysis of the chicken reproductive tract transcriptome.

Analysis of the chicken reproductive tract transcriptome is important in comparative biology for analysis of reproductive tract development and evolution. In addition, molecular analysis of the reproductive tract is important for identification of genes affecting fertility in the poultry industry. We sampled the chicken reproductive tract (ovary, oviduct, and testis) transcriptome, generating 5,328 expressed sequence tags that assembled into 4,518 contigs. We identified 475 contigs with no match in the current expressed sequence tag databases or in GenBank. The novel contigs included 31 with no match to the current assembly of the chicken genome, 119 representing spliced transcripts, and 309 that were unspliced. More detailed molecular characterization of the 428 novel contigs present in the assembly will be important to gene discovery and annotation of the chicken and other vertebrate genomes.

Aging↗

Large-scale generation and analysis of expressed sequence tags from porcine ovary.

One method to identify the factors that control ovarian function is to characterize the genes that are expressed in ovary. In the present study, cDNA libraries from fetal, neonatal, and prepubertal porcine ovaries, pubertal ovaries on different days of the estrous cycle (Days 0 [follicle], 5, and 12 [follicle and corpus luteum]), and follicles isolated from weaned sows (diameter, 2, 4, 6, and 8 mm) were constructed and sequenced. A total of 22 176 cDNAs were sequenced, of which 15 613 were of sufficient quality for clustering. Clustering of cDNAs resulted in 8507 contigs, 6294 (74%) of which were comprised of a single sequence. Sixty-eight percent of the contigs had consensus sequences that were homologous to existing Tentative Consensus (TC) sequences or mature transcripts (ET) in The Institute for Genomic Research Porcine Gene Index. The consensus sequences were classified according to the Gene Ontology Index. Most cDNA-encoded proteins were components of the nucleus, ribosome, or mitochondrion. The proteins primarily functioned in binding, catalysis, and transport. Nearly 75% of the proteins were involved in metabolism and cell growth and/or maintenance. Analysis of the cDNA frequency across different libraries demonstrated differential gene expression within different-size follicles, between follicles and corpora lutea, and across developmental time-points. The expression of selected genes (analyzed by ribonuclease protection assay and Northern blotting) was consistent with the frequency of their respective cDNA in the individual libraries. This porcine ovary unigene set will be useful for identifying factors and mechanisms controlling ovarian follicular development in a variety of species.

Animals↗

Organization of the Bacillus subtilis 168 chromosome between kdg and the attachment site of the SP beta prophage: use of Long Accurate PCR and yeast artificial chromosomes for sequencing.

Within the Bacillus subtilis genome sequencing project, the region between lysA and ilvA was assigned to our laboratory. In this report we present the sequence of the last 36 kb of this region, between the kdg operon and the attachment site of the SP beta prophage. A two-step strategy was used for the sequencing. In the first step, total chromosomal DNA was cloned in phage M13-based vectors and the clones carrying inserts from the target region were identified by hybridization with a cognate yeast artificial chromosome (YAC) from our collection. Sequencing of the clones allowed us to establish a number of contigs. In the second step the contigs were mapped by Long Accurate (LA) PCR and the remaining gaps closed by sequencing of the PCR products. The level of sequence inaccuracy due to LA PCR errors appeared to be about 1 in 10,000, which does not affect significantly the final sequence quality. This two-step strategy is efficient and we suggest that it can be applied to sequencing of longer chromosomal regions. The 36 kb sequence contains 38 coding sequences (CDSs), 19 of which encode unknown proteins. Seven genetic loci already mapped in this region, xpt, metB, ilvA, ilvD, thyB, dfrA and degR were identified. Eleven CDSs were found to display significant similarities to known proteins from the data banks, suggesting possible functions for some of the novel genes: cspD may encode a cold shock protein; bcsA, the first bacterial homologue of chalcone synthase; exol, a 5' to 3' exonuclease, similar to that of DNA polymerase I of Escherichia coli; and bsaA, a stress-response-associated protein. The protein encoded by yplP has homology with the transcriptional NifA-like regulators. The arrangement of the genes relative to possible promoters and terminators suggests 19 potential transcription units.

Amino Acid Sequence↗

FHY1: a phytochrome A-specific signal transducer.

Phytochromes are plant photoreceptors that regulate plant growth and development with respect to the light environment. Following the initial light-perception event, the phytochromes initiate a signal-transduction process that eventually results in alterations in cellular behavior, including gene expression. Here we describe the molecular cloning and functional characterization of Arabidopsis FHY1. FHY1 encodes a product (FHY1) that specifically transduces signals downstream of the far-red (FR) light-responsive phytochrome A (PHYA) photoreceptor. We show that FHY1 is a novel light-regulated protein that accumulates in dark (D)-grown but not in FR-grown hypocotyl cells. In addition, FHY1 transcript levels are regulated by light, and by the product of FHY3, another gene implicated in FR signaling. These observations indicate that FHY1 function is both FR-signal transducing and FR-signal regulated, suggesting a negative feedback regulation of FHY1 function. Seedlings homozygous for loss-of-function fhy1 alleles are partially blind to FR, whereas seedlings overexpressing FHY1 exhibit increased responses to FR, but not to white (WL) or red (R) light. The increased FR-responses conferred by overexpression of FHY1 are abolished in a PHYA-deficient mutant background, showing that FHY1 requires a signal from PHYA for function, and cannot modulate growth independently of PHYA.

Alleles↗

Sequence diversity and genomic organization of vomeronasal receptor genes in the mouse.

The vomeronasal system of mice is thought to be specialized in the detection of pheromones. Two multigene families have been identified that encode proteins with seven putative transmembrane domains and that are expressed selectively in subsets of neurons of the vomeronasal organ. The products of these vomeronasal receptor (Vr) genes are regarded as candidate pheromone receptors. Little is known about their genomic organization and sequence diversity, and only five sequences of mouse V1r coding regions are publicly available. Here, we have begun to characterize systematically the V1r repertoire in the mouse. We isolated 107 bacterial artificial chromosomes (BACs) containing V1r genes from a 129 mouse library. Hybridization experiments indicate that at least 107 V1r-like sequences reside on these BACs. We assembled most of the BACs into six contigs, of which one major contig and one minor contig were characterized in detail. The major contig is 630-860 kb long, encompasses a cluster of 21-48 V1r genes, and contains marker D6Mit227. Sequencing of the coding regions was facilitated by the absence of introns. We determined the sequence of the coding region of 25 possibly functional V1r genes and seven pseudogenes. The functional V1rs can be arranged into three groups; V1rs of one group are novel and substantially divergent from the other V1rs. The genomic and sequence information described here should be useful in defining the biological function of these receptors.

Animals↗

Annotating large genomes with exact word matches.

We have developed a tool for rapidly determining the number of exact matches of any word within large, internally repetitive genomes or sets of genomes. Thus we can readily annotate any sequence, including the entire human genome, with the counts of its constituent words. We create a Burrows-Wheeler transform of the genome, which together with auxiliary data structures facilitating counting, can reside in about one gigabyte of RAM. Our original interest was motivated by oligonucleotide probe design, and we describe a general protocol for defining unique hybridization probes. But our method also has applications for the analysis of genome structure and assembly. We demonstrate the identification of chromosome-specific repeats, and outline a general procedure for finding undiscovered repeats. We also illustrate the changing contents of the human genome assemblies by comparing the annotations built from different genome freezes.

Algorithms↗