Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

Empirical analysis of transcriptional activity in the Arabidopsis genome.

Functional analysis of a genome requires accurate gene structure information and a complete gene inventory. A dual experimental strategy was used to verify and correct the initial genome sequence annotation of the reference plant Arabidopsis. Sequencing full-length cDNAs and hybridizations using RNA populations from various tissues to a set of high-density oligonucleotide arrays spanning the entire genome allowed the accurate annotation of thousands of gene structures. We identified 5817 novel transcription units, including a substantial amount of antisense gene transcription, and 40 genes within the genetically defined centromeres. This approach resulted in completion of approximately 30% of the Arabidopsis ORFeome as a resource for global functional experimentation of the plant proteome.

Arabidopsis↗

A statistical approach for array CGH data analysis.

BACKGROUND: Microarray-CGH experiments are used to detect and map chromosomal imbalances, by hybridizing targets of genomic DNA from a test and a reference sample to sequences immobilized on a slide. These probes are genomic DNA sequences (BACs) that are mapped on the genome. The signal has a spatial coherence that can be handled by specific statistical tools. Segmentation methods seem to be a natural framework for this purpose. A CGH profile can be viewed as a succession of segments that represent homogeneous regions in the genome whose BACs share the same relative copy number on average. We model a CGH profile by a random Gaussian process whose distribution parameters are affected by abrupt changes at unknown coordinates. Two major problems arise: to determine which parameters are affected by the abrupt changes (the mean and the variance, or the mean only), and the selection of the number of segments in the profile. RESULTS: We demonstrate that existing methods for estimating the number of segments are not well adapted in the case of array CGH data, and we propose an adaptive criterion that detects previously mapped chromosomal aberrations. The performances of this method are discussed based on simulations and publicly available data sets. Then we discuss the choice of modeling for array CGH data and show that the model with a homogeneous variance is adapted to this context. CONCLUSIONS: Array CGH data analysis is an emerging field that needs appropriate statistical tools. Process segmentation and model selection provide a theoretical framework that allows precise biological interpretations. Adaptive methods for model selection give promising results concerning the estimation of the number of altered regions on the genome.

Algorithms↗

HLA-G gene polymorphism segregation within CEPH reference families.

HLA-G, a nonclassical HLA class I antigen, presents tissue-restricted expression on human trophoblasts and may play an important role in immune tolerance of mother-versus-fetus. In this work we have demonstrated extensive HLA-G genomic polymorphism within three CEPH reference families, by PCR-SSCP analysis and direct sequencing. Among six unrelated parents we assigned eight HLA-G alleles, seven of which are new. We observed the segregation of HLA-G alleles of heterozygous parents among their offspring that matched the segregation of the HLA class I haplotypes. Only one of the mutations observed was found to be nonsynonymous indicating low polymorphism of the HLA-G molecule.

Alleles↗

DNA fingerprinting for differentiation of field isolates from reference vaccine strains of Pasteurella multocida in turkeys.

The genomes from field isolates of Pasteurella multocida in turkeys and those of P multocida reference CU and M9 vaccine strains were analyzed and compared after cleavage with restriction endonucleases. The electrophoretic profiles obtained with DNA fragments from field isolates and vaccine strains of the same serotype were characteristic and reproducible. These features indicated the existence of differences among the isolates of the same serotype that cannot currently be detected, using available serotyping methods. However, several field isolates had electrophoretic profiles similar to those of either CU or M9 vaccine strain. It was concluded that restriction endonuclease analysis of DNA genomes from P multocida isolated from turkeys provides the information for differentiation of field isolates from vaccine strains of the same serotype.

Animals↗

Genome conservation in isolates of Leptospira interrogans.

Reference strains for each of the 23 serogroups of Leptospira interrogans yielded different pulsed-field gel electrophoresis patterns of NotI digestion products. This was also the case for the 14 serovars belonging to serogroup Icterohaemorrhagiae (with one exception). The NotI restriction patterns of 45 clinical leptospiral isolates belonging to serovar icterohaemorrhagiae were analyzed and compared with those of type strains. No differences were observed between isolates from countries of different continents, namely, France, French Guiana, New Caledonia, and Tahiti. The pattern was indistinguishable from that of the reference strain of serovar icterohaemorrhagiae.

Biological Evolution↗

Trypanosoma cruzi genome project: biological characteristics and molecular typing of clone CL Brener.

Clone CL Brener is the reference organism used in the Trypanosoma cruzi Genome Project. CL Brener was obtained by cloning procedures from bloodstream trypomastigotes isolated from mice infected with the CL strain. The doubling time of CL Brener epimastigotes cultured at 28 degrees C in liver infusion-tryptose (LIT) medium is 58 +/- 13 h. Differentiation to metacyclic forms is induced by incubation of epimastigotes in LIT-20% Grace's medium. Metacyclics give very low parasitemia in mice, contrary to what is observed for blood forms which promote 100% mortality of the animals with inocula of 5 x 10(3) parasites. CL Brener blood forms are highly susceptible to nifurtimox, benznidazole and ketoconazole. Allopurinol is inefficient in the treatment of mice experimental infection. The clone infects mammalian cultured cells and performs the complete intracellular cycle at 33 and 37 degrees C. The molecular typing of CL Brener has been done by isoenzymatic profiles; sequencing of a 24S alpha ribosomal RNA gene domain and by schizodeme, randomly amplified polymorphic DNA and DNA fingerprinting analyses. For each typing approach the patterns obtained do not change after prolonged parasite subcultivation in LIT medium (up to 100 generations). The stability of the molecular karyotype of the clone was also confirmed.

Animals↗

Maintaining collections of mutants for plant functional genomics.

As the plant genomics era progresses and post-genomic functional research rapidly expands, varied genetic resources of unprecedented power and scope are being developed. Partially by the mandate of public funding, these resources are being shared via stock centers and private laboratories. The successful initiation of any new research requires that advantage be taken of these stocks. Information on most plant genomic resources can be obtained through simple yet powerful. Web searches, and ordering mechanisms are linked to the information. Hence, locating and obtaining materials is rapid and simple. Currently, available genomic resources are described, and references, links for Web data, and ordering information are also included.

Arabidopsis↗

Transposon-like properties of the major, long repetitive sequence family in the genome of Physarum polycephalum.

A family of long, highly-repetitive sequences, referred to previously as ;HpaII-repeats', dominates the genome of the eukaryotic slime mould Physarum polycephalum. These sequences are found exclusively in scrambled clusters. They account for about one-half of the total complement of repetitive DNA in Physarum, and represent the major sequence component found in hypermethylated, 20-50 kb segments of Physarum genomic DNA that fail to be cleaved using the restriction endonuclease HpaII. The structure of this abundant repetitive element was investigated by analysing cloned segments derived from the hypermethylated genomic DNA compartment. We show that the ;HpaII-repeat' forms part of a larger repetitive DNA structure, approximately 8.6 kb in length, with several structural features in common with recognised eukaryotic transposable genetic elements. Scrambled clusters of the sequence probably arise as a result of transposition-like events, during which the element preferentially recombines in either orientation with target sites located in other copies of the same repeated sequence. The target sites for transposition/recombination are not related in sequence but in all cases studied they are potentially capable of promoting the formation of small ;cruciforms' or ;Z-DNA' structures which might be recognised during the recombination process.

Journal Article↗

Ten millennia of purifying selection on HLA-B27 reveals an ancient epidemic-scale burden of spondyloarthritis in West Eurasia.

HLA-B27 exemplifies an evolutionary trade-off between protection against infection and susceptibility to inflammatory disease. To investigate its long-term population history, this study examined three HLA-B27-tagging variants-rs116488202, rs4349859, and rs116666910-in present-day populations from UK biobank and 1000 Genomes Project, and in 15,836 ancient West Eurasian individuals. The estimated frequency of HLA-B27 reached 49.0% approximately 8500 years before present, then declined progressively to 3.9% in the present-day reference population. Comparison with genome-wide association study (GWAS) data for ankylosing spondylitis (AS) showed that HLA-B27-linked alleles conferring increased disease risk had negative selection coefficients, indicating sustained selection against HLA-B27 over the past 10,000 years. The decline coincided with major Holocene changes in settlement and subsistence patterns, microbial exposure, and enteric infection, which may have increased the inflammatory costs of HLA-B27. Because previous paleopathological studies have largely been limited to identifying advanced skeletal manifestations of AS, this ancient genetic study may provide currently the most sensitive population-level record of an otherwise largely undetectable, epidemic-scale disease burden in antiquity. These findings support the hypothesis that HLA-B27-associated spondyloarthritis was sufficiently prevalent and severe to influence human evolution in prehistoric West Eurasia.

Humans↗

Genome mining reveals an architecturally expanded pyoluteorin-associated biosynthetic gene cluster and a divergent flavin-dependent halogenase-like sequence in deep-sea Pseudomonas Aeruginosa from the Gulf of Guinea.

BACKGROUND: Marine deep-sea environments harbour microorganisms with extraordinary biosynthetic potential, yet their secondary metabolite repertoires remain largely uncharacterised. RESULTS: This study reports the isolation, phenotypic characterisation, and whole-genome analysis of Pseudomonas aeruginosa strain E1, recovered from deep Atlantic seawater (Gulf of Guinea, ~2500 m depth), which exhibits antifungal activity against multidrug-resistant Candida parapsilosis. Three presumptive P. aeruginosa isolates (E1, E17, and E44) showed > 99% 16S rRNA gene sequence identity to P. aeruginosa reference sequences, while whole-genome dDDH analysis of strain E1 yielded 95.2% (95% CI: 93.6-96.4%; formula d4) relative to the P. aeruginosa type strain DSM 50071ᵀ (= ATCC 10145ᵀ), supporting its species-level assignment. Antifungal screening and PCR-based detection of flavin-dependent halogenase genes identified strain E1 as the primary candidate for genomic investigation. Illumina whole-genome sequencing produced a 6.33 Mb draft genome assembly (113 contigs, 5862 protein-coding genes, 66.4% GC content). Genome mining with antiSMASH 8.0 identified 27 biosynthetic gene clusters (BGCs) spanning nonribosomal peptide synthetase (NRPS), polyketide synthase (PKS), phenazine, terpene, and metallophore pathways. Region 7.1 of strain E1 harbours a predicted 50.8 kb pyoluteorin-associated BGC, comprising 34 genes, substantially larger than its terrestrial counterpart (~ 22 kb, ~ 17 genes), and featuring nine transport genes and three regulatory elements. Phylogenetic analysis resolved three halogenase genes: ctg7_146 showed 98.7% amino acid identity to PltA, and ctg7_149 showed 99.2% amino acid identity to PltM, supporting their annotation as PltA-like and PltM-like components of the predicted pyoluteorin biosynthetic pathway. Among the characterised reference enzymes included in this analysis, ctg7_143 showed the highest amino acid identity to PltM from P. fluorescens Pf-5. However, the identity remained low at approximately 30.4%, supporting its placement as a divergent FDH-like sequence rather than a close PltM orthologue. CONCLUSION: This study provides the first comprehensive genomic characterisation of a pyoluteorin-BGC-harbouring marine P. aeruginosa strain, demonstrating conservation of the core biosynthetic machinery alongside an expanded transport architecture and a divergent FDH-like sequence that may represent a candidate for future biochemical investigation. These findings expand current knowledge of FDH-like sequence diversity in deep-sea bacteria and support further investigation of Gulf of Guinea microorganisms as a potential source of biosynthetic and enzymatic diversity.

Multigene Family↗

Diversifying selection in human papillomavirus type 16 lineages based on complete genome analyses.

Human papillomavirus type 16 (HPV16) is the primary etiological agent of cervical cancer, the second most common cancer in women worldwide. Complete genomes of 12 isolates representing the major lineages of HPV16 were cloned and sequenced from cervicovaginal cells. The sequence variations within the open reading frames (ORFs) and noncoding regions were identified and compared with the HPV16R reference sequence. This whole-genome approach gives us unprecedented precision in detailing sequence-level changes that are under selection on a whole-viral-genome scale. Of 7,908 base pair nucleotide positions, 313 (4.0%) were variable. Within the 2,452 amino acids (aa) comprising 8 ORFs, 243 (9.9%) amino acid positions were variable. In order to investigate the molecular evolution of HPV16 variants, maximum likelihood models of codon substitution were used to identify lineages and amino acid sites under selective pressure. Five codon sites in the E5 (aa 48, 65) and E6 (aa 10, 14, 83) ORFs were demonstrated to be under diversifying selective pressure. The E5 ORF had the overall highest nonsynonymous/synonymous substitution rate (omega) ratio (M3 = 0.7965). The E2 gene had the next-highest omega ratio (M3 = 0.5611); however, no specific codons were under positive selection. These data indicate that the E6 and E5 ORFs are evolving under positive Darwinian selection and have done so in a relatively short time period. Whether response to selective pressure upon the E5 and E6 ORFs contributes to the biological success of HPV16, its specific biological niche, and/or its oncogenic potential remains to be established.

Adult↗

Genomic variations in echovirus 30 persistent isolates recovered from a chronically infected immunodeficient child and comparison with the reference strain.

Seven sequential isolates of echovirus type 30 (EV30) were recovered over 22 months from a child with severe combined immune deficiency syndrome. The nucleotide sequences of the 5' halves of the genomes (4,400 nucleotides) of the first (S1) and last (S7) isolates were determined and compared with that of the EV30 Bastianni reference strain, also determined in this study. In genome regions P1 and P2, 101 variations were identified between the two isolates. Synonymous differences far outnumbered nonsynonymous differences. Amino acid changes affected both capsid and nonstructural polypeptides (particularly 2B). The VP1 nucleotide sequences of the seven isolates were determined to analyze genome evolution during the chronic infection. In the phylogenetic tree, the seven isolates were directly related to the prototype strain in an individual monophyletic group, strongly suggesting that the chronic infection in the child arose from a single persistent EV30 isolate. Four lineages were observed in the persistent isolates. Isolates S2, S4, S5, and S6 were close relatives of one another, whereas isolates S1 and S3 formed individual lineages. Isolate S7, distantly related to all other isolates, formed the fourth lineage. These findings suggest the quasispecies nature of the genomes of the seven sequential EV30 isolates. Grouping of persistent isolates on the basis of replicative capacities was consistent with phylogenetic relationships. Overall, the results indicate that genetically related EV30 variants with different replicative capacities coexisted in a carrier state, probably in the gastrointestinal tract, during the infection of the child.

5' Untranslated Regions↗

The ORFanage: an ORFan database.

As each newly sequenced genome contains a significant number of protein-coding ORFs that are species-, family- or lineage-specific, many interesting questions arise about the evolution and role of these ORFs and of the genomes they are part of. We refer to these poorly conserved ORFs as singleton or paralogous ORFans if they are unique to one genome, or as orthologous ORFans if they appear only in a family of closely related organisms and have no homolog in other genomes. In order to study and classify ORFans we have constructed the ORFanage, an ORFan database. This database consists of the predicted ORFs in fully sequenced microbial genomes, and enables searching for the three types of ORFans in any subset of the genomes chosen by the user. The ORFanage could help in choosing interesting targets for further genomic and evolutionary studies. The ORFanage is accessible via http://www.bioinformatics.buffalo. edu/ORFanage.

Computational Biology↗

Insertions, deletions, and single-nucleotide polymorphisms at rare restriction enzyme sites enhance discriminatory power of polymorphic amplified typing sequences, a novel strain typing system for Escherichia coli O157:H7.

Polymorphic amplified typing sequences (PATS) for Escherichia coli O157:H7 (O157) was previously based on indels containing XbaI restriction enzyme sites occurring in O-island sequences of the O157 genome. This strain-typing system, referred to as XbaI-based PATS, typed every O157 isolate tested in a reproducible, rapid, straightforward, and easy-to-interpret manner and had technical advantages over pulsed-field gel electrophoresis (PFGE). However, the system was less discriminatory than PFGE and was unable to differentiate fully between unrelated isolates. To overcome this drawback, we enhanced PATS by using another infrequently cutting restriction enzyme, AvrII (also known as BlnI), to identify additional polymorphic regions that could increase the discriminatory ability of PATS typing. Referred to as AvrII-based PATS, the system identified seven new polymorphic regions in the O157 genome. Unlike XbaI, polymorphisms involving AvrII sites were caused by both indels and single-nucleotide polymorphisms occurring in O-island and backbone sequences of the O157 genome. AvrII-based PATS by itself provided poor discrimination of the O157 isolates tested. However, when primer pairs amplifying the seven polymorphic AvrII sites were combined with those amplifying the eight polymorphic XbaI sites (combined PATS), the discriminatory power of PATS was enhanced. Combined PATS matched related O157 isolates better than PFGE while differentiating between unrelated isolates. PATS typed every O157 isolate tested and directly targeted polymorphic sequences responsible for differences in the restriction digest patterns of O157 genomic DNA, utilizing PCR rather than relying on gel electrophoresis. This enabled PATS to resolve the ambiguity in PFGE typing, including that arising from the "more distantly related" and "untypeable" profiles.

Bacterial Typing Techniques↗

Identification of putative homology between horse microsatellite flanking sequences and cross-species ESTs, mRNAs and genomic sequences.

In this study the flanking sequences of 1534 horse microsatellites were used in a BLAST search to identify putative human-horse homologies. BLAST searches revealed 129 flanking sequences with significant blastn matches [alignment scores (S) > or = 60 and sum probability values (E) < or = 3.0E-6], also, 25 of these produced significant blastx matches. To provide a reference point in the human genome the flanking sequences with matches were subjected to a BLAT search of the University of California Santa Cruz (UCSC) human genome assembly (July 2003 freeze). Eighty-three of the flanking sequences showed high similarity to sequence of known or putative human genes and the remaining 46 demonstrated high similarity to human intragenic regions. Interestingly, 87 of the microsatellites showed conservation of the tandem repeat in addition to flanking regions. Overall, 41 of the microsatellites had been mapped in the horse and of these 37 localized to the expected syntenic location. The other four did not and represent new putative regions of human-horse synteny. The results of this study contribute 79 new putative human-horse homologies, increasing the density of markers on the human-horse comparative map.

Animals↗

High resolution physical map of porcine chromosome 7 QTL region and comparative mapping of this region among vertebrate genomes.

BACKGROUND: On porcine chromosome 7, the region surrounding the Major Histocompatibility Complex (MHC) contains several Quantitative Trait Loci (QTL) influencing many traits including growth, back fat thickness and carcass composition. Previous studies highlighted that a fragment of approximately 3.7 Mb is located within the Swine Leucocyte Antigen (SLA) complex. Internal rearrangements of this fragment were suggested, and partial contigs had been built, but further characterization of this region and identification of all human chromosomal fragments orthologous to this porcine fragment had to be carried out. RESULTS: A whole physical map of the region was constructed by integrating Radiation Hybrid (RH) mapping, BAC fingerprinting data of the INRA BAC library and anchoring BAC end sequences on the human genome. 17 genes and 2 reference microsatellites were ordered on the high resolution IMNpRH212000rad Radiation Hybrid panel. A 1000:1 framework map covering 550 cR12000 was established and a complete contig of the region was developed. New micro rearrangements were highlighted between the porcine and human genomes. A bovine RH map was also developed in this region by mapping 16 genes. Comparison of the organization of this region in pig, cattle, human, mouse, dog and chicken genomes revealed that 1) the translocation of the fragment described previously is observed only on the bovine and porcine genomes and 2) the new internal micro rearrangements are specific of the porcine genome. CONCLUSION: We estimate that the region contains several rearrangements and covers 5.2 Mb of the porcine genome. The study of this complete BAC contig showed that human chromosomal fragments homologs of this heavily rearranged QTL region are all located in the region of HSA6 that surrounds the centromere. This work allows us to define a list of all candidate genes that could explain these QTL effects.

Animals↗

The Universal Declaration on the Human Genome and Human Rights.

Since 1985, UNESCO studies ethical questions arising in genetics. In 1992, I established the International Bioethics Committee at UNESCO with the mission to draft the Universal Declaration on the Human Genome and Human Rights, which was adopted by UNESCO in 1997 and the United Nations in 1998. The Declaration relates the human genome with human dignity, deals with the rights of the persons concerned by human genome research and provides a reference legal framework for both stimulating the ethical debate and the harmonization of the law worldwide, favouring useful developments that respect human dignity.

Codes of Ethics↗