Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Genomic imprinting: implications for human disease.

Genomic imprinting refers to an epigenetic marking of genes that results in monoallelic expression. This parent-of-origin dependent phenomenon is a notable exception to the laws of Mendelian genetics. Imprinted genes are intricately involved in fetal and behavioral development. Consequently, abnormal expression of these genes results in numerous human genetic disorders including carcinogenesis. This paper reviews genomic imprinting and its role in human disease. Additional information about imprinted genes can be found on the Genomic Imprinting Website at http://www.geneimprint.com.

Angelman Syndrome↗

A refined molecular karyotype for the reference strain of the Trypanosoma cruzi genome project (clone CL Brener) by assignment of chromosome markers.

We present a useful refinement of the molecular karyotype of clone CL Brener, the reference clone of the Trypanosoma cruzi Genome Project. The assignment of 210 genetic markers (142 expressed sequence tags (ESTs), seven cDNAs, 32 protein-coding genes, eight sequence tagged sites (STSs), 21 repetitive sequences) to the chromosomal bands separated by pulsed field gel electrophoresis (PFGE) identified 61 chromosome-specific markers, two size-polymorphic chromosomes and seven linkage groups. Fourteen new repetitive elements were isolated in this work and mapped to the chromosomal bands. We found that at least ten repetitive elements can be mapped to each chromosomal band, which may render the whole genome sequence assembly a difficult task. To construct the integrated map of chromosomal band XX, we used yeast artificial chromosome (YAC) overlapping clones and a variety of probes (i.e. known gene sequences, ESTs, STSs generated from the YAC ends). The total length covered by the YAC contig was approximately 1.3 Mb, covering 37% of the entire chromosome. We found some degree of polymorphism among YACs derived from band XX. These results are in agreement with data from phylogenetic analysis of T. cruzi which suggest that clone CL Brener is a hybrid genotype [Mol. Biochem. Parasitol. 92 (1998) 253; Proc. Natl. Acad. Sci. USA 98 (2001) 7396]. The physical map of the chromosomal bands, together with the isolation of specific chromosomal markers, will contribute in the global effort to sequence the nuclear genome of this parasite.

Animals↗

[Construction and expression of Neisseria surface protein (nspA) of Neisseria gonorrhoeae in Escherichia coli].

OBJECTIVE: To construct neisseria surface protein (NspA) recombinants of Neisseria gonorrhoeae from a reference strain and express this protein in E. coli. METHODS: The fragments of NspA gene of Neisseria gonorrhoeae was amplified by PCR from the reference strain genomic DNA and cloned into expression vector pET-30c (+) to get the pET-NspA recombinants. The recombinants were verified with restrictive endonuclease digestion and sequence analysis. The verified recombinant was transformed into E. coli BL21 (DE3). After inducing with IPTG, the expressed NspA protein was analyzed by SDS-PAGE and Western Blot. RESULTS: The pET-NspA expression recombinants for the reference strain of Neisseria gonorrhoeae were successfully constructed and the induced recombinant NspA protein was observed. CONCLUSION: The successful expression of the Neisseria gonorrhoeae NspA protein will be very helpful for the further research of its antigenicity and immunological activity, and for the construction of preventive vaccines on Neisseria gonorrhoeae infection.

Bacterial Outer Membrane Proteins↗

GCRP: Integrated Global Chicken Reference Panel from 11,951 Chicken Genomes.

Chickens are a crucial source of protein for humans and a popular model animal for bird research. Despite the emergence of imputation as a reliable genotyping strategy for large populations, the lack of a high-quality chicken reference panel has hindered progress in chicken genome research. To address this, here we introduce the first phase of the 100K Global Chicken Reference Panel (100K GCRP). Currently, two panels are available: a comprehensive mix panel (CMP) for domestication diversity research and a commercial breed panel (CBP) for breeding broilers specifically. Evaluation of genotype imputation quality showed that CMP had the highest imputation accuracy compared to imputation using existing chicken panels in Animal-SNPAtlas and Animal Genotype Imputation Database (AGIDB), whereas CBP performed stably in the imputation of commercial populations. Additionally, we found that genome-wide association studies using GCRP-imputed data, whether on simulated or real phenotypes, exhibited greater statistical power. In conclusion, our study indicates that the GCRP effectively fills the gap in high-quality reference panels for chickens, providing an effective imputation platform for future genetic and breeding research. The project includes 11,951 samples and provides services for various applications on its website at http://farmrefpanel.com/GCRP/#/.

Animals↗

Evidence for large inversion polymorphisms in the human genome from HapMap data.

Knowledge about structural variation in the human genome has grown tremendously in the past few years. However, inversions represent a class of structural variation that remains difficult to detect. We present a statistical method to identify large inversion polymorphisms using unusual Linkage Disequilibrium (LD) patterns from high-density SNP data. The method is designed to detect chromosomal segments that are inverted (in a majority of the chromosomes) in a population with respect to the reference human genome sequence. We demonstrate the power of this method to detect such inversion polymorphisms through simulations done using the HapMap data. Application of this method to the data from the first phase of the International HapMap project resulted in 176 candidate inversions ranging from 200 kb to several megabases in length. Our predicted inversions include an 800-kb polymorphic inversion at 7p22, a 1.1-Mb inversion at 16p12, and a novel 1.2-Mb inversion on chromosome 10 that is supported by the presence of two discordant fosmids. Analysis of the genomic sequence around inversion breakpoints showed that 11 predicted inversions are flanked by pairs of highly homologous repeats in the inverted orientation. In addition, for three candidate inversions, the inverted orientation is represented in the Celera genome assembly. Although the power of our method to detect inversions is restricted because of inherently noisy LD patterns in population data, inversions predicted by our method represent strong candidates for experimental validation and analysis.

Biometry↗

A reference cross DNA panel for zebrafish (Danio rerio) anchored with simple sequence length polymorphisms.

The ultimate informativeness of the zebrafish mutations described in this issue will rest in part on the ability to clone these genes. However, the genetic infrastructure required for the positional cloning in zebrafish is still in its infancy. Here we report a reference cross panel of DNA, consisting of 520 F2 progeny (1040 meioses) that has been anchored to a zebrafish genetic linkage map by 102 simple sequence length polymorphisms. This reference cross DNA provides: (1) a panel of DNA from the cross that was used to construct the genetic linkage map, upon which polymorphic gene(s) and genetic markers can be mapped; (2) a fine order mapping tool, with a maximum resolution of 0.1 cM; and (3) a foundation for the development of a physical map (an ordered array of clones each containing a known portion of the genome). This reference cross DNA will serve as a resource enabling investigators to relate genes or genetic markers directly to a single genetic linkage map and avoid the problem of integrating different maps with different genetic markers, as must be currently done when using randomly amplified polymorphic DNA markers, or as has occurred with human genetic linkage maps.

Alleles↗

A high-resolution 6.0-megabase transcript map of the type 2 diabetes susceptibility region on human chromosome 20.

Recent linkage studies and association analyses indicate the presence of at least one type 2 diabetes susceptibility gene in human chromosome region 20q12-q13.1. We have constructed a high-resolution 6.0-megabase (Mb) transcript map of this interval using two parallel, complementary strategies to construct the map. We assembled a series of bacterial artificial chromosome (BAC) contigs from 56 overlapping BAC clones, using STS/marker screening of 42 genes, 43 ESTs, 38 STSs, 22 polymorphic, and 3 BAC end sequence markers. We performed map assembly with GraphMap, a software program that uses a greedy path searching algorithm, supplemented with local heuristics. We anchored the resulting BAC contigs and oriented them within a yeast artificial chromosome (YAC) scaffold by observing the retention patterns of shared markers in a panel of 21 YAC clones. Concurrently, we assembled a sequence-based map from genomic sequence data released by the Human Genome Project, using a seed-and-walk approach. The map currently provides near-continuous coverage between SGC32867 and WI-17676 ( approximately 6.0 Mb). EST database searches and genomic sequence alignments of ESTs, mRNAs, and UniGene clusters enabled the annotation of the sequence interval with experimentally confirmed and putative transcripts. We have begun to systematically evaluate candidate genes and novel ESTs within the transcript map framework. So far, however, we have found no statistically significant evidence of functional allelic variants associated with type 2 diabetes. The combination of the BAC transcript map, YAC-to-BAC scaffold, and reference Human Genome Project sequence provides a powerful integrated resource for future genomic analysis of this region.

Base Composition↗

[Impaired epigenetic gene activity regulation and human diseases].

The epigenetic (i.e. heritable states that are mediated by changes in DNA other than nucleotide sequence) mechanisms of regulation of gene expression have been recently the focus of intensive studies. Genomic imprinting refers to the epigenetic gene marking that results in monoallelic expression. The epigenetic mechanism of imprinting is based on the gamete-specific methylation of some mammalian genes, which restricts their expression on one of the parental chromosomes. The imprinted genes control fetal and placental development, cell proliferation and adult behavior. Changes in the normal imprinting patterns give rise to numerous genetic diseases, including cancer. Examining the molecular processes that mediate these methylation genome changes will give use a great insight into the mechanisms of regulation of gene activity and into the etiology of some human genetic diseases.

Alleles↗

A variable region gene subfamily encoding T cell receptor beta-chains is selectively conserved among mammals.

Studies of murine T cell receptor genes indicate that the VT beta repertoire, in contrast to the large VT alpha, VH, and VK repertoires, is limited to approximately 21 genes. Large differences between the various VT beta sequences allow classification into distinct subfamilies consisting of one or few members. The VT beta sequence of a gene, RTB92, expressed in a rabbit T cell line is most closely related to a VT beta gene expressed in the human cell line Molt-4. Genomic blots with RTB92 VT beta region used as probe indicated conservation of related VT beta gene subfamilies in man and pig, but no related sequences were seen in rat, mouse, or hamster. The relationship among these VT beta genes mirrors similarities and differences in MHC genes of the compared species and suggests coordinate evolution of these functionally related gene families. The constant region of RTB92 was shown to be rabbit CT beta 2 by reference to genomic clones. Comparisons with constant region sequences of human and murine CT beta 1 and CT beta 2 chains reveal no similarities in amino acid or nucleic acid sequence in the coding region characteristic of the respective isotypes. Intraspecies homologies between CT beta 1 and CT beta 2 sequences were much greater than those between CT beta 1 or CT beta 2 from other species. By contrast, comparisons of JT beta and 3' untranslated region sequences showed significant interspecies conservation of sequences in the beta 1 and in the beta 2 regions.

Amino Acid Sequence↗

The genome of Eimeria spp., with special reference to Eimeria tenella--a coccidium from the chicken.

Eimeria spp. contain at least four genomes. The nuclear genome is best studied in the avian species Eimeria tenella and comprises about 60 Mbp DNA contained within ca. 14 chromosomes; other avian and lupine species appear to possess a nuclear genome of similar size. In addition, sequence data and hybridisation studies have provided direct evidence for extrachromosomal mitochondrial and plastid DNA genomes, and double-stranded RNA segments have also been described. The unique phenotype of "precocious" development that characterises some selected lines of Eimeria spp. not only provides the basis for the first generation of live attenuated vaccines, but offers a significant entrée into studies on the regulation of an apicomplexan life-cycle. With a view to identifying loci implicated in the trait of precocious development, a genetic linkage map of the genome of E. tenella is being constructed in this laboratory from analyses of the inheritance of over 400 polymorphic DNA markers in the progeny of a cross between complementary drug-resistant and precocious parents. Other projects that impinge directly or indirectly on the genome and/or genetics of Eimeria spp. are currently in progress in several laboratories, and include the derivation of expressed sequence tag data and the development of ancillary technologies such as transfection techniques. No large-scale genomic DNA sequencing projects have been reported.

Animals↗

Trypanosoma cruzi 5S rRNA arrays define five groups and indicate the geographic origins of an ancestor of the heterozygous hybrids.

Isolates of the etiological agent of Chagas disease, Trypanosoma cruzi, have been subdivided into six subgroups referred to as discrete typing units. The subgroups are related through two distinct hybridisation events: representatives of homozygous discrete typing units I and IIb fused to form discrete typing units IIa and IIc, whose homozygous genotypes have features of both ancestral types; a second fusion between strains of homozygous discrete typing units IIb and IIc created the heterozygous hybrid strains discrete typing units IId and IIe. The intergenic region of the tandemly repeated 5S rRNA array displays four variant sequence classes, allowing the discrimination of five discrete typing units. The genome project reference strain, CL Brener, is a hybrid discrete typing unit IIe strain that contains both discrete typing unit IIb and IIc classes of 5S rRNA repeats in distinct arrays present on different chromosomes. The CL Brener discrete typing unit IIb-type array contains approximately 193 repeated units, of which about one-third contain a 129 bp sequence that replaces a majority of the 5S rRNA sequence. The 129 bp 'invader' sequence was detected within the arrays of all hybrid discrete typing unit IId and IIe strains and in a subset of discrete typing unit IIb strains. This array invader replaces the internal promoter elements conserved in 5S rRNA. The discrete typing unit IIb Esmeraldo strain contains approximately 135 repeats and shows a region of homology to the array invader in the 5' flank of the array, but no evidence of the invading sequence element within the array. A survey of additional discrete typing unit IIb strains revealed a split within the subgroup, in which some strains contained invaded arrays and others were homogeneous for the 5S rRNA. The putative discrete typing unit IIb ancestor of the hybrid discrete typing units IId and IIe more closely resembles the extant Bolivian/Chilean IIb isolates than the Brazilian IIb isolates based on the correlation with the array invader.

Animals↗

DNA methylation in leprosy-associated bacteria: Mycobacterium leprae and Corynebacterium tuberculostearicum.

The DNAs of two kinds of microorganisms from human leprosy lesion, Mycobacterium leprae and Corynebacterium tuberculostearicum (also known as "leprosy-derived corynebacterium" or LDC), have been analysed and compared with the genomes of reference bacteria of the CMN group (genera Corynebacterium, Mycobacterium and Nocardia). The guanine-plus-cytosine content (% GC) of DNA was determined by a double-labelling procedure, which is unaffected by the presence of modified and unusual bases (that alter both buoyant density and mid-melting-point determinations). Accordingly, the DNAs of seven LDC strains had GC values of 54-56 mol %, and that of armadillo-grown M. leprae a value of 54.8 +/- 0.9 mol %. Restriction patterns disclosed no methylated cytosine in the DNA sequences CCGG, GGCC, AGCT and GATC of either LDC or M. leprae DNA. N6-methyl adenine was present in the sequence GATC of all LDC strains, but was missing from the genomes of all others CMN organisms analysed, including M. leprae. By HPLC analysis of LDC-DNA hydrolysates, it was found that N6-methyladenine amounted to 1.8% of total DNA adenine, and was present exclusively within GATC sequences, which appeared all to be methylated. It is concluded that LDC represent a group of corynebacteria endowed with high genetic homogeneity and a unique restriction pattern, whereby their genome is easily distinguished from that of M. leprae, which has a similar base composition.

Adenine↗

Simple repetitive (GAA)n loci in the human genome.

In order to investigate the organization (and inheritance) of simple tandem (GAA)n repeats in the human genome, different restriction enzymes were employed for DNA digestion followed by separation of the resulting fragments by agarose gel electrophoresis. Frequently cutting enzymes (4 bp recognition sites) revealed highly complex multilocus banding patterns after conventional horizontal submarine gels. Fragments larger than 25 kb, resulting from digestions with rarely cutting enzymes (6 bp recognition sites), were separated by pulsed-field gel electrophoresis (PFGE). Hybridizations were carried out directly in the gel matrix. Nearly all of the enzymes produce at least one predominant signal band after hybridization with (GAA)6. The frequency of (GAA)n stretches was estimated on human chromosome 4 by probing a respective cosmid library. In comparison with other simple di-, tri- and tetranucleotide repeats, (GAA)n stretches appear underrepresented. Hybridization of a fetal human brain cDNA library indicated very few expressed (GAA)n repeats. These data are discussed with particular reference to genomic organization and other simple repetitive trinucleotide stretches.

Chromosome Mapping↗

ESTAnnotator: A tool for high throughput EST annotation.

In high throughput sequence analysis, it is often necessary to combine the results of contemporary bioinformatics tools, because no individual tool alone computes all the requested information. ESTAnnotator is a tool for the high throughput annotation of expressed sequence tags (ESTs) by automatically running a collection of bioinformatics applications. In the first step, a quality check is performed and repeats, vector parts and low quality sequences are masked. Then successive steps of database searching and EST clustering are performed. Already known transcripts present within mRNA and genomic DNA reference databases are identified. Subsequently, tools for the clustering of anonymous ESTs, and for further database searches at the protein level, are applied. Finally, the outputs of each individual tool are gathered and the relevant results presented in a descriptive summary. ESTAnnotator was already successfully applied for the systematic identification and characterisation of novel human genes involved in cartilage/bone formation, growth, differentiation and homeostasis. ESTAnnotator is available at http://genome.dkfz-heidelberg.de, contact: genome@dkfz.de.

Cartilage↗

Genetic diversity of bradyrhizobial populations from diverse geographic origins that nodulate Lupinus spp. and Ornithopus spp.

The genetic diversity of 45 bradyrhizobial isolates that nodulate several Lupinus and Ornithopus species in different geographic locations was investigated by 16S rDNA PCR-RFLP and sequence analysis, 16S-23S rDNA intergenic spacer (IGS) PCR-RFLP analysis, and ERIC-PCR genomic fingerprinting. Reference strains of Bradyrhizobium japonicum, B. liaoningense and B. elkanii and some Canarian isolates from endemic woody legumes in the tribe Genisteae were also included. The 16S rDNA-RFLP analysis resolved 9 genotypes of lupin isolates, a group of fourteen isolates presented restriction-genotypes identical or very similar to B. japonicum, while another two main groups of isolates (69%) presented genotypes that clearly separated them from the reference species of soybean. 16S rDNA sequencing of representative strains largely agreed with restriction analysis, except for a group of six isolates, and showed that all the lupin isolates are relatives of B. japonicum, but different lineages were observed. The 16S-23S IGS-RFLP analysis showed a high resolution level, resolving 19 distinct genotypes among 30 strains analysed, and so demonstrating the heterogeneity of the 16S-RFLP groups. ERIC-PCR fingerprint analysis showed an enormous genetic diversity producing a different pattern for each but two of the isolates. Phylogeny of nodC gene was independent from the 16S rRNA phylogeny, and showed a tight relationship in the symbiotic region of the lupin isolates with isolates from Canarian genistoid woody legumes, and in concordance, cross-nodulation was found. We conclude that Lupinus is a promiscuous host legume that is nodulated by rhizobia with very different chromosomal genotypes, which could even belong to several species of Bradyrhizobium. No correlation among genomic background, original host plant and geographic location was found, so, different chromosomal genotypes could be detected at a single site and in a same plant species, on the contrary, an identical genotype was detected in very different geographical locations and plants.

Bacterial Proteins↗

Comparative karyotyping as a tool for genome structure analysis of Trypanosoma cruzi.

As a part of the Trypanosoma cruzi genome project, 239 genetic markers were hybridised to PFGE separated DNA from T. cruzi, in order to determine the number and size of chromosomes and to aid the assembly of the genome sequence. We used three strains, T. cruzi IIe CL Brener (the genome project reference strain) and two T. cruzi I strains, Sylvio X10/7 and CAI/72, to perform a comparative study of their karyotypes and to determine marker linkage. A densitometry analysis of the separations estimated the total chromosome numbers to be 55 in CL Brener and 57 in the two other strains. In all, 45 markers hybridised to single chromosomal bands and 103 markers to two bands in CL Brener, while the number of markers in Sylvio X10/7 and CAI/72 were 102/68 and 61/105, respectively. Size differences between homologous chromosomes were often large, up to 1900 kb (173%). The average difference was 36% for CL Brener and 23.5% for the T. cruzi I strains. Larger differences in CL Brener are consistent with a recent hybrid origin. Forty markers distributed into 15 linkage groups were found to identify specific chromosomes or chromosomes pairs. While the same markers are generally linked in all three strains, the sizes of the chromosomes vary extensively, indicating large chromosomal rearrangements. These data provide valuable information for the finishing of the CL Brener genome sequence.

Animals↗

Regulation of genomic imprinting by gametic and embryonic processes.

Parental genomic imprinting refers to the phenomenon by which alleles behave differently depending on the sex of the parent from which they are inherited. In the case of the murine transgene RSVIgmyc, imprinting is manifest in two ways: differential DNA methylation and differential expression. In inbred FVB/N mice, a transgene inherited from a male parent is undermethylated and expressed; a transgene inherited from the female parent is overmethylated and silent. Using a series of RSVIgmyc constructs and transgenic mice, we show that the imprinting of this transgene requires a cis-acting signal that is principally derived from the repeat sequences that make up the 3' portion of the murine immunoglobulin alpha heavy-chain switch region. Such imprinting is relatively independent of the site of transgene insertion but is influenced by the structure of the transgene itself. Imprinting is also modulated by genetic background. Detailed studies indicate that the paternal allele is undermethylated and expressed in inbred FVB/N mice and in heterozygous F1 FVB/N/C57Bl/6J mice but is overmethylated and silent in inbred C57Bl/6J mice. Consequently, the FVB/N genome appears to carry alleles of modulating genes that dominantly block methylation and permit expression of the paternally imprinted transgene. Furthermore, our results suggest that overmethylation is the default status of both parental alleles and that the paternal allele can be marked in trans by polymorphic factors that act in postblastocyst embryos.

Age Factors↗

Codon usage comparison of novel genes in clinical isolates of Haemophilus influenzae.

A similarity statistic for codon usage was developed and used to compare novel gene sequences found in clinical isolates of Haemophilus influenzae with a reference set of 80 prokaryotic, eukaryotic and viral genomes. These analyses were performed to obtain an indication as to whether individual genes were Haemophilus-like in nature, or if they probably had more recently entered the H.influenzae gene pool via horizontal gene transfer from other species. The average and SD values were calculated for the similarity statistics from a study of the set of all genes in the H.influenzae Rd reference genome that encoded proteins of 100 amino acids or longer. Approximately 80% of Rd genes gave a statistic indicating that they were most like other Rd genes. Genes displaying codon usage statistics >1 SD above this range were either considered part of the highly expressed group of H.influenzae genes, or were considered of foreign origin. An alternative determinant for identifying genes of foreign origin was when the similarity statistics produced a value that was much closer to a non-H.influenzae reference organism than to any of the Haemophilus species contained in the reference set. Approximately 65% of the novel sequences identified in the H.influenzae clinical isolates displayed codon usages most similar to Haemophilus sp. The remaining novel sequences produced similarity statistics closer to one of the other reference genomes thereby suggesting that these sequences may have entered the H.influenzae gene pool more recently via horizontal transfer.

Base Sequence↗