Access to every trial dataset is crucial.
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Large-scale expression data, such as that generated by hybridization to microarrays, is potentially a rich source of information on gene function and regulation. By clustering genes according to their expression profiles, groups of genes involved in the same pathways or sharing common regulatory mechanisms may be identified. Publicly-available EST collections are a largely unexplored source of expression data. We previously used a sample of rice ESTs to generate 'digital expression profiles' by counting the frequency of tags for different genes sequenced from different cDNA libraries. A simple statistical test was used to associate genes or cDNA libraries having similar expression profiles. Here we further validate this approach using larger samples of ESTs from the UniGene projects (clustered human, mouse and rat ESTs). Our results show that genes clustered on the basis of expression profile may represent genes implicated in similar pathways or coding for different subunits of multi-component enzyme complexes. In addition we suggest that comparison of clusters from different species, may be useful for confirmation or prediction of orthologs.
Age-related macular degeneration (AMD) is a multifactorial disorder affecting the visual system with a high prevalence among the elderly population but with no effective therapy available at present. To better understand the pathogenesis of this disorder, the identification of the genetic factors and the determination of their contribution to AMD is needed. Towards this goal, we are pursuing a strategy that makes use of the EST data processed in the UniGene database and aims at the generation of a comprehensive catalogue of genes preferentially active in the human retina. Subsequently, these genes will be systematically assessed in AMD. We performed a retina EST sampling and obtained a total of 673 clusters containing only retina ESTs as well as 568 clusters with at least 30% of the ESTs in each cluster originating from retina cDNA libraries. Of these, 180 representative EST clusters with varying retina and non-retina EST contents were analyzed for their in vitro expression. This approach identified 39 transcripts with retina-specific expression. One of these genes (C18orf2) mapping to chromosome 18 was further characterized. Multiple C18orf2 transcripts display a complex pattern of differential splicing in the human retina. The various isoforms encode hypothetical polypeptides with no homologies to known proteins or protein motifs.
Patients diagnosed with a standard clinical method (subject to misclassification error) are often combined with patients diagnosed with a gold-standard method (with zero or very small misclassification error) in family-based studies of complex disease. For example, non-autopsied patients (NAP) are often included along with autopsy-proven (AP) patients in family-based studies of complex diseases, such as Alzheimer's disease (AD). Theoretical and simulation studies suggest that certain misclassification errors can result in severe reduction of power in genetic linkage and association analyses and that phenotype (or diagnostic) error can produce misleading results. Morton's test for heterogeneity can identify genomic regions where error may have led to loss in power. We applied this test to pedigree data from the NIMH Alzheimer's Disease Genetics Initiative Database separated into AP and NAP pedigrees. Morton's test identified one highly significant region of heterogeneity on chromosome 2. The source of the heterogeneity was due to significant indication of linkage in the AP pedigrees at position 109 cM (p value = 6.68 x 10(-5)) with no indication in the NAP pedigrees. Furthermore, Morton's test showed no evidence for heterogeneity on chromosome 19 in early-onset pedigrees that showed highly significant evidence for linkage in other published reports. These results suggest that supplementing linkage analysis with Morton's test can be usefully applied to genetic data sets that have AP and NAP samples, or other sample mixtures that include a 'gold standard' subgroup with reduced error rate, to increase power to detect linkage in the presence of diagnostic misclassification.
The Central Clydeside Conurbation (CCC) has relatively high mortality rates. This paper examines whether it also has relatively high rates of ill health, using data from three cohorts (aged 15, 35 and 55 in 1987/88) in the West of Scotland. Comparisons on a range of self-reported physical and mental health indicators, anthropometric measures, blood pressure, and respiratory function were made with comparable age groups in ten British or Scottish national studies. The older two cohorts in the CCC exhibited relatively high rates of longstanding and limiting longstanding illness and the youngest cohort had relatively poor psychosocial health, compared to their age peers elsewhere. Fewer differences were found in blood pressure, anthropometric measures or respiratory function although older CCC residents were slightly shorter than in Britain as a whole and had slightly poorer respiratory function. Central Clydesiders in the late 1980s were generally in poorer health than those of the same sex and similar age elsewhere in the UK, but the extent of the disadvantage varied across different dimensions of health, and was not as marked as some stereotypes of the West of Scotland would suggest.
BACKGROUND: The availability of increasing amounts of sequence data from completely sequenced genomes boosts the development of new computational methods for automated genome annotation and comparative genomics. Therefore, there is a need for tools that facilitate the visualization of raw data and results produced by bioinformatics analysis, providing new means for interactive genome exploration. Visual inspection can be used as a basis to assess the quality of various analysis algorithms and to aid in-depth genomic studies. RESULTS: GeneViTo is a JAVA-based computer application that serves as a workbench for genome-wide analysis through visual interaction. The application deals with various experimental information concerning both DNA and protein sequences (derived from public sequence databases or proprietary data sources) and meta-data obtained by various prediction algorithms, classification schemes or user-defined features. Interaction with a Graphical User Interface (GUI) allows easy extraction of genomic and proteomic data referring to the sequence itself, sequence features, or general structural and functional features. Emphasis is laid on the potential comparison between annotation and prediction data in order to offer a supplement to the provided information, especially in cases of "poor" annotation, or an evaluation of available predictions. Moreover, desired information can be output in high quality JPEG image files for further elaboration and scientific use. A compilation of properly formatted GeneViTo input data for demonstration is available to interested readers for two completely sequenced prokaryotes, Chlamydia trachomatis and Methanococcus jannaschii. CONCLUSIONS: GeneViTo offers an inspectional view of genomic functional elements, concerning data stemming both from database annotation and analysis tools for an overall analysis of existing genomes. The application is compatible with Linux or Windows ME-2000-XP operating systems, provided that the appropriate Java Runtime Environment is already installed in the system.
BACKGROUND: The learning of global genetic regulatory networks from expression data is a severely under-constrained problem that is aided by reducing the dimensionality of the search space by means of clustering genes into putatively co-regulated groups, as opposed to those that are simply co-expressed. Be cause genes may be co-regulated only across a subset of all observed experimental conditions, biclustering (clustering of genes and conditions) is more appropriate than standard clustering. Co-regulated genes are also often functionally (physically, spatially, genetically, and/or evolutionarily) associated, and such a priori known or pre-computed associations can provide support for appropriately grouping genes. One important association is the presence of one or more common cis-regulatory motifs. In organisms where these motifs are not known, their de novo detection, integrated into the clustering algorithm, can help to guide the process towards more biologically parsimonious solutions. RESULTS: We have developed an algorithm, cMonkey, that detects putative co-regulated gene groupings by integrating the biclustering of gene expression data and various functional associations with the de novo detection of sequence motifs. CONCLUSION: We have applied this procedure to the archaeon Halobacterium NRC-1, as part of our efforts to decipher its regulatory network. In addition, we used cMonkey on public data for three organisms in the other two domains of life: Helicobacter pylori, Saccharomyces cerevisiae, and Escherichia coli. The biclusters detected by cMonkey both recapitulated known biology and enabled novel predictions (some for Halobacterium were subsequently confirmed in the laboratory). For example, it identified the bacteriorhodopsin regulon, assigned additional genes to this regulon with apparently unrelated function, and detected its known promoter motif. We have performed a thorough comparison of cMonkey results against other clustering methods, and find that cMonkey biclusters are more parsimonious with all available evidence for co-regulation.
BACKGROUND: Analyses of genetic data at the level of haplotypes provide increased accuracy and power to infer genotype-phenotype correlations and evolutionary history of a locus. However, empirical determination of haplotypes is expensive and laborious. Therefore, several methods of inferring haplotypes from unphased genotypic data have been proposed, but it is unclear how accurate each of the methods is or which methods are superior. The accuracy of some of the leading methods of computational haplotype inference (PL-EM, Phase, SNPHAP, Haplotyper) are compared using a large set of 308 empirically determined haplotypes based on 15 SNPs, among which 36 haplotypes were observed to occur. This study presents several advantages over many previous comparisons of haplotype inference methods: a large number of subjects are included, the number of known haplotypes is much smaller than the number of chromosomes surveyed, a range in values of linkage disequilibrium, presence of rare SNP alleles, and considerable dispersion in the frequencies of haplotypes. RESULTS: In contrast to some previous comparisons of haplotype inference methods, there was very little difference in the accuracy of the various methods in terms of either assignment of haplotypes to individuals or estimation of haplotype frequencies. Although none of the methods inferred all of the known haplotypes, the assignment of haplotypes to subjects was about 90% correct for individuals heterozygous for up to three SNPs and was about 80% correct for up to five heterozygous sites. All of the methods identified every haplotype with a frequency above 1%, and none assigned a frequency above 1% to an incorrect haplotype. CONCLUSIONS: All of the methods of haplotype inference have high accuracy and one can have confidence in inferences made by any one of the methods. The ability to identify even rare (>/= 1%) haplotypes is reassuring for efforts to identify haplotypes that contribute to disease in a significant proportion of a population. Assignment of haplotypes is relatively accurate among subjects heterozygous for up to 5 sites, and this might be the largest number of SNPs for which one should define haplotype blocks or have confidence in haplotype assignments.
Recent studies have suggested that a high-density single nucleotide polymorphism (SNP) marker set could provide equivalent or even superior information compared with currently used microsatellite (STR) marker sets for gene mapping by linkage. The focus of this study was to compare results obtained from linkage analyses involving extended pedigrees with STR and single-nucleotide polymorphism (SNP) marker sets. We also wanted to compare the performance of current linkage programs in the presence of high marker density and extended pedigree structures. One replicate of the Genetic Analysis Workshop 14 (GAW14) simulated extended pedigrees (n = 50) from New York City was analyzed to identify the major gene D2. Four marker sets with varying information content and density on chromosome 3 (STR [7.5 cM]; SNP [3 cM, 1 cM, 0.3 cM]) were analyzed to detect two traits, the original affection status, and a redefined trait more closely correlated with D2. Multipoint parametric and nonparametric linkage analyses (NPL) were performed using programs GENEHUNTER, MERLIN, SIMWALK2, and S.A.G.E. SIBPAL. Our results suggested that the densest SNP map (0.3 cM) had the greatest power to detect linkage for the original trait (genetic heterogeneity), with the highest LOD score/NPL score and mapping precision. However, no significant improvement in linkage signals was observed with the densest SNP map compared with STR or SNP-1 cM maps for the redefined affection status (genetic homogeneity), possibly due to the extremely high information contents for all maps. Finally, our results suggested that each linkage program had limitations in handling the large, complex pedigrees as well as a high-density SNP marker set.
BACKGROUND: We have developed a simulation-based approach to the analysis of shared homozygous chromosomal segments and have applied it to data on allele sharing among alcoholics in a single Collaborative Study on the Genetics of Alcoholism pedigree. Our assessment of sharing involved the use of a single-nucleotide polymorphism (SNP) marker map provided by Affymetrix. RESULTS: All 11 affected individuals in the selected pedigree shared 2 copies of an allele at 4 adjacent SNPs in a region on chromosome 5. Via simulation, we determined that the probability that such sharing is caused by mere chance is less than 0.0000001. After correcting for undocumented inbreeding, this probability rose to 0.0016. The probability that the shared segment emanates from a single ancestor and is unrelated to the affection status is less than 0.0000001 in the corrected pedigree. Haplotype association analysis and a search for a protective locus using unaffected individuals yielded no significant results. CONCLUSION: Homozygosity mapping results on chromosome 5 provide suggestive evidence of the region's role as one that may harbor a genetic determinant of alcoholism. Furthermore, the probabilities of chance homozygous allele sharing for the original and for the inbreeding-corrected pedigree provide insight into the impact that inbreeding can have on such calculations.
We combined the results of whole-genome linkage and association analyses to determine which markers were most strongly associated with Kofendrerd Personality Disorder. Using replicate 1 from the Genetic Analysis Workshop 14 Aipotu, Karangar, Danacaa, and New York City simulated populations, we determined that several markers showed significant linkage and association with disease status. We used both SNP and microsatellite markers to determine patterns and chromosomal regions of markers. Three consistently associated markers were C01R0050, C03R0280, and C10R0882. Using generalized linear mixed models, we modelled the effect of the three predefined phenotypic categories on disease status and concluded that the phenotypes defining the "anxiety-related" category best predicted the outcome.
Genome scans using dense single-nucleotide polymorphism (SNP) data have recently become a reality. It is thought that the increase in information content for linkage analysis as a result of the denser scans will help refine previously identified linkage regions and possibly identify new regions not identifiable using the sparser, microsatellite scans. In the context of the dense SNP scans, it is also possible to consider association strategies to provide even more information about potential regions of interest. To circumvent the multiple-testing issues inherent in association analysis, we use a recently developed strategy, implemented in PBAT, which screens the data to identify the optimal SNPs for testing, without biasing the nominal significance level. We compare the results from the PBAT analysis to that of quantitative linkage analysis on chromosome 4 using the Collaborative Study on the Genetics of Alcoholism data, as released through Genetic Analysis Workshop 14.
Many decisions about genome sequencing projects are directed by perceived gaps in the tree of life, or towards model organisms. With the goal of a better understanding of biology through the lens of evolution, however, there are additional genomes that are worth sequencing. One such rationale for whole-genome sequencing is discussed here, along with other important strategies for understanding the phenotypic divergence of species.
BACKGROUND: The spread of OXA-48-like carbapenemases represents a major public health challenge. Although previous studies have investigated OXA-48-like carbapenemases risk factors, nosocomial dissemination, and plasmid dynamics, an integrated plasmid-centered framework combining complete plasmid mining, transmission-unit analysis, phylogenetic reconstruction, and machine learning-based risk assessment remains limited. METHODS: We systematically collected 747 complete plasmid sequences carrying blaOXA-48-like genes from the NCBI database, establishing the largest collections of complete plasmid sequences to date. Using an integrative framework of population genomics, phylogenetic dating, and machine learning, this study aimed to characterize the dissemination patterns, plasmid replicon diversity, transmission units, mobile genetic elements, co-resistance profiles, and risk classification of these plasmid. RESULTS: Plasmids carrying blaOXA-48-like genes were detected across 50 countries on six continents, with blaOXA-48 predominating in Europe, blaOXA-181 in South Asia, and blaOXA-232 largely in Asia. IncL and ColKP3/IncX3 replicons, together with Tn1999.2 and other MGEs, were central drivers of plasmid maintenance and spread. Sixteen transmission units were defined, with AA068_Cluster3 estimated to have originated in the Netherlands around 2005 before expanding to Europe, the Middle East, Asia, and North America. Co-resistance analyses revealed frequent modules involving aminoglycoside and quinolone resistance, with qnrS1 and aph(3'')-Ib most prevalent. Notably, high-risk transposon structures were often identified in non-clinical environments, underscoring their cross-ecological transmission potential. Machine learning-based classification models showed good internal performance for predefined composite-risk categories, with plasmid mobility, clinical/non-clinical source composition, and host background contributing to the classification results. CONCLUSIONS: This study provides a large-scale plasmid-centered genomic analysis of publicly available complete plasmid sequences carrying blaOXA-48-like genes, integrating transmission-unit inference, phylogeographic reconstruction, mobile genetic element and co-resistance profiling, and composite genomic risk stratification. This gene-centered framework may support future One Health-oriented antimicrobial resistance surveillance and prioritization of plasmids with higher dissemination and resistance potential.