Beyond release: the equitable use of genomic information.
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
BACKGROUND: Searching for approximate patterns in large promoter sequences frequently produces an exceedingly high numbers of results. Our aim was to exploit biological knowledge for definition of a sheltered search space and of appropriate search parameters, in order to develop a method for identification of a tractable number of sequence motifs. RESULTS: Novel software (COOP) was developed for extraction of sequence motifs, based on clustering of exact or approximate patterns according to the frequency of their overlapping occurrences. Genomic sequences of 1 Kb upstream of 91 genes differentially expressed and/or encoding proteins with relevant function in adult human retina were analyzed. Methodology and results were tested by analysing 1,000 groups of putatively unrelated sequences, randomly selected among 17,156 human gene promoters. When applied to a sample of human promoters, the method identified 279 putative motifs frequently occurring in retina promoters sequences. Most of them are localized in the proximal portion of promoters, less variable in central region than in lateral regions and similar to known regulatory sequences. COOP software and reference manual are freely available upon request to the Authors. CONCLUSION: The approach described in this paper seems effective for identifying a tractable number of sequence motifs with putative regulatory role.
Toll-like receptors recognize pathogen-associated molecular patterns (PAMPs) and TLR5 is the pathogen recognition receptor (PRR) for bacterial flagellin. Patients carrying a R392 stop polymorphism display an inflammatory phenotype and increased susceptibility to pneumonia caused by the flagellated bacteria Legionella pneumophila. While this suggests that TLR5 mutations may be clinically relevant, functional data are not available for the majority of the other TLR5 polymorphisms. We have characterized all known single nucleotide polymorphisms (SNPs) of TLR5 for their functional relevance upon stimulation in transiently transfected CHO-K1 cells. Among the 13 missense SNPs of TLR5 reported in the human genetic databases, three SNPs (c.1174C>T, p.R392X; c.2081A>G, p.D694G; and c.2464C>T, p.L822F) were found to be functionally relevant in transiently transfected CHO-K1 cells. The prevalences of these functionally relevant SNPs in our investigation were 11.9 %, 0 %, and 0 %, in healthy donors. The p.D694G and p.L822F SNPs are of low frequency in the Caucasian population though further investigations of the common p.R392X variant alone or of functional relevant TLR5 SNPs in combination with other TLR SNPs will elucidate their possible role on disease susceptibility in humans and may facilitate clinical diagnosis.
OBJECTIVE: We evaluated the medical-sociological implications of parental perception of risk and decision-making choices for prenatally ascertained choroid plexus cysts (CPCs) between two obstetric populations with similar clinical situations. METHODS: The Wayne State University (WSU) Reproductive Genetics database and the Madigan Army Medical Center (MAMC) experience were reviewed to compare the rates of aneuploidy and invasive testing for cases with CPC. Aneuploidy rates were compared between those with isolated CPC, CPC with advanced maternal age (AMA), and CPC associated with multiple anomalies. RESULTS: 186 cases were identified in the WSU cohort, of whom 27 (15%) declined invasive fetal testing. In the remaining 159 cases, aneuploidy was present in 2/132 (1.5%) isolated CPCs, 3/11 (27%) CPCs with AMA, and 15/16 (93%) CPCs with multiple anomalies. 107 cases were identified in the MAMC cohort, of whom 99 (92%) declined invasive fetal testing. No cases of aneuploidy were found in the 3/12 AMA cases or 5/95 non-AMA cases who underwent amniocentesis. CONCLUSIONS: The 2 cases of aneuploidy with isolated CPC cannot be ignored, and provide an estimated attributable risk of at least 0.8%, a higher risk than 38 years of age. However, the parental sociologic context may be as important as the genetic-prognostic risk for decision-making.
We evaluated the medical-sociological implications of parental perception of risk and decision-making choices for prenatally ascertained choroid plexus cysts (CPC) between two obstetric populations. The Wayne State University (WSU) Reproductive Genetics database and the Madigan Army Medical Center (MAMC) experience were reviewed to compare the rates of aneuploidy and invasive testing for cases with CPC. Aneuploidy rates were compared between those with isolated CPC, CPC with advanced maternal age (AMA), and CPC associated with multiple anomalies. In the WSU cohort 186 cases were identified, of whom 27 (15%) declined invasive fetal testing. In the remaining 159 cases, aneuploidy was present in 2/132 (1.5%) isolated CPC, 3/11 (27%) CPC with AMA, and 15/16 (93%) CPC with multiple anomalies. In the MAMC cohort 107 cases were identified, of whom 99 (92%) declined invasive fetal testing. No aneuploidy cases were found in the 3/12 AMA cases or 5/95 non-AMA cases that underwent amniocentesis. The two cases of aneuploidy with isolated CPC cannot be ignored, and provide an estimated attributable risk of at least 0.8%, a higher risk than 38 years of age. However, the parental sociologic context may be as important for decision-making as the genetic-prognostic risk.
BACKGROUND: Cis-regulatory modules are combinations of regulatory elements occurring in close proximity to each other that control the spatial and temporal expression of genes. The ability to identify them in a genome-wide manner depends on the availability of accurate models and of search methods able to detect putative regulatory elements with enhanced sensitivity and specificity. RESULTS: We describe the implementation of a search method for putative transcription factor binding sites (TFBSs) based on hidden Markov models built from alignments of known sites. We built 1,079 models of TFBSs using experimentally determined sequence alignments of sites provided by the TRANSFAC and JASPAR databases and used them to scan sequences of the human, mouse, fly, worm and yeast genomes. In several cases tested the method identified correctly experimentally characterized sites, with better specificity and sensitivity than other similar computational methods. Moreover, a large-scale comparison using synthetic data showed that in the majority of cases our method performed significantly better than a nucleotide weight matrix-based method. CONCLUSION: The search engine, available at http://mapper.chip.org, allows the identification, visualization and selection of putative TFBSs occurring in the promoter or other regions of a gene from the human, mouse, fly, worm and yeast genomes. In addition it allows the user to upload a sequence to query and to build a model by supplying a multiple sequence alignment of binding sites for a transcription factor of interest. Due to its extensive database of models, powerful search engine and flexible interface, MAPPER represents an effective resource for the large-scale computational analysis of transcriptional regulation.
On the basis of comparison of the cytochrome b gene nucleotide sequences from genetic databases, the possible phylogenetic relationships of mitochondrial DNA (mtDNA) among all major lineages of Salmoninae (Brachymystax, Parahucho, Salvelinus, Salmo, Parasalmo, and Oncorhynchus) were examined. Three different phylogenetic methods (UPGMA, NJ, and ML) yielded phylogenetic trees of essentially the same topology: (((Brachymystax, Parahucho), Salvelinus, Salmo), (Parasalmo, Oncorhynchus)). The results obtained using the maximum parsimony method were less clear. Apparently, the divergence of the main salmonid lineages occurred during a relatively short time period; hence, the number of synapomorphs marking the order of their divergence was extremely low. This may account for the relative failure to use the maximum parsimony method of phylogenetic reconstruction. The problem of concordance of mtDNA and species phylogenetic schemes is discussed. Their discrepancy in salmonids may be caused by interspecific introgressive hybridization.
The MEGADATS relational database system has many useful applications in the field of medical genetics. Some of these applications include storage, retrieval, and display of pedigree information; retrieval of sets of individuals, sibships, or families who meet given criteria; storage of necessary information for mailing lists, clinic data, etc; and combination of pedigree information and genotype information into the format needed for linkage analysis packages.
As the amount of biological data grows, so does the need for biologists to store and access this information in central repositories in a free and unambiguous manner. The European Bioinformatics Institute (EBI) hosts six core databases, which store information on DNA sequences (EMBL-Bank), protein sequences (SWISS-PROT and TrEMBL), protein structure (MSD), whole genomes (Ensembl) and gene expression (ArrayExpress). But just as a cell would be useless if it couldn't transcribe DNA or translate RNA, our resources would be compromised if each existed in isolation. We have therefore developed a range of tools that not only facilitate the deposition and retrieval of biological information, but also allow users to carry out searches that reflect the interconnectedness of biological information. The EBI's databases and tools are all available on our website at www.ebi.ac.uk.
The emergence of new technologies from the genomics revolution will transform the potential application of biomarkers to assess how pollutants impact people, animals, and ecosystems. Genetic databases provide a huge resource from which candidate molecular biomarkers can be identified and, subsequently, exploited to address these issues. However, a major challenge is to link these novel molecular indices to ecologically relevant whole-organism life-cycle traits (such as reproduction and growth). Such a functional link is provided by annetocin, previously characterized as a member of the vasopressin/oxytocin superfamily of neuropeptides. It is expressed in annelid worms within the neurons of the central nervous system and has been shown to be involved in the induction of egg-laying behavior. This paper outlines the validation of annetocin as a novel biomarker of reproductive fitness in the earthworm Eisenia fetida. The design of primer pairs targeted toward oligochaete annetocin has facilitated the isolation of a full-length annetocin cDNA from this species. Optimization of a real-time quantitative PCR procedure exploiting the fluorescent DNA-binding molecule, Sybr Green, has allowed the measurement of annetocin transcript levels over a range covering six orders of magnitude. Using this approach, gene expression was measured in earthworms exposed to soils polluted with high concentrations of zinc and lead. Traditional growth and reproductive indices, including cocoon production, were also recorded and related to the molecular parameter. The future use of annetocin as a molecular genetic biomarker in terrestrial ecotoxicology is discussed.
The use of hypervariable tandem repeat loci for population genetic studies, genetic analysis of inherited disease and individual identification purposes requires establishment of a genetic database for each reference population. In the present study we have analysed variability at five tandem repeat loci (D1S80, D17S5, 3'-hvr/apoB, F8vWF and D6S89)in a representative sample (88 to 156 individuals of greek ancestry), using polymerase chain reaction amplification. Between nine and 19 alleles were resolved throughout the five polymorphic loci. Heterozygosity indices for these loci in the greek population ranged from 0.68 to 0.85. Allele frequencies follow a bimodal discrimination (pd) and allelic diversity (h) values ranged from 0.84 to 0.94 and 0.85 to 0.91, respectively, and indicated that these loci are highly informative and can be used for population studies, forensic purposes and parentage and family testing. Comparison of observed and expected genotype frequencies by the conventional chi-square test indicated conformity to Hardy-Weinberg predictions.
Viruses are intracellular parasites that use many cellular pathways during their replication. Large DNA viruses, such as herpesviruses, have captured a repertoire of cellular genes to block or mimic host immune responses, apoptosis regulation, and cell-cycle control mechanisms. We have conducted a systematic search for all homologs of herpesvirus proteins in the human genome using position-specific scoring matrices representing herpesvirus protein sequence domains, and pair-wise sequence comparisons. The analysis shows that approximately 13% of the herpesvirus proteins have clear sequence similarity to products of the human genome. Different human herpesviruses vary in their numbers of human homologs, indicating distinct rates of gene acquisition in different lineages. Our analysis has identified new families of herpesvirus/human homologs from viruses including human herpesvirus 5 (human cytomegalovirus; HCMV) and human herpesvirus 8 (Kaposi's sarcoma-associated herpesvirus; KSHV), which may play important roles in host-virus interactions.
Based on biomedical literature databases, we tried a first step for constructing a gene expression "data warehouse" specific to human colorectal cancer (CRC). Results of genome-wide transcriptomic research were available from 12 studies, using various technologies, namely, SAGE, cDNA and oligonucleotide arrays, and adaptor-tagged amplification. Three studies analyzed CRC cell lines and nine studies of human samples. The total number of patients was 144. Out of 982 up- or down-regulated genes, 863 (88%) were found to be differentially expressed in a single study, 88 in two studies, 22 in three studies, 7 in four studies, and only 2 genes in six studies. Eight large-scale proteomics studies were published in CRC, using 2-D-, SDS- or free-flow electrophoresis, involving only 11 patients. Out of 408 differentially expressed proteins, 339 (83%) were found to be differentially expressed only in a single study, 16 in three studies, 10 in four studies, 3 in five, and 1 in eight studies. Confirmation at proteome level of results obtained with large-scale transcriptomics studies was possible in 25%. This proportion was higher (67%) for reproducing proteome results using transcriptomics technologies. Obviously, reproducibility and overlapping between published gene expression results at proteome and transcriptome level are low in human CRC. Thus, the development of standardized processes for collecting samples, storing, retrieving, and querying gene expression data obtained with different technologies is of central importance in translational research.
Helicobacter pylori is naturally competent for DNA transformation, but the mechanism by which transformation occurs is not known. For Haemophilus influenzae, dprA is required for transformation by chromosomal but not plasmid DNA, and the complete genomic sequence of H. pylori 26695 revealed a dprA homolog (HP0333). Examination of genetic databases indicates that DprA homologs are present in a wide variety of bacterial species. To examine whether HP0333 has a function similar to dprA of H. influenzae, HP0333, present in each of 11 strains studied, was disrupted in two H. pylori isolates. For both mutants, the frequency of transformation by H. pylori chromosomal DNA was markedly reduced, but not eliminated, compared to their wild-type parental strains. Mutation of HP0333 also resulted in a marked decrease in transformation frequency by a shuttle plasmid (pHP1), which differs from the phenotype described in H. influenzae. Complementation of the mutant with HP0333 inserted in trans in the chromosomal ureAB locus completely restored the frequency of transformation to that of the wild-type strain. Thus, while dprA is required for high-frequency transformation, transformation also may occur independently of DprA. The presence of DprA homologs in bacteria known not to be naturally competent suggests a broad function in DNA processing.
BACKGROUND: Microarray devices permit a genome-scale evaluation of gene function. This technology has catalyzed biomedical research and development in recent years. As many important diseases can be traced down to the gene level, a long-standing research problem is to identify specific gene expression patterns linking to metabolic characteristics that contribute to disease development and progression. The microarray approach offers an expedited solution to this problem. However, it has posed a challenging issue to recognize disease-related genes expression patterns embedded in the microarray data. In selecting a small set of biologically significant genes for classifier design, the nature of high data dimensionality inherent in this problem creates substantial amount of uncertainty. RESULTS: Here we present a model for probability analysis of selected genes in order to determine their importance. Our contribution is that we show how to derive the P value of each selected gene in multiple gene selection trials based on different combinations of data samples and how to conduct a reliability analysis accordingly. The importance of a gene is indicated by its associated P value in that a smaller value implies higher information content from information theory. On the microarray data concerning the subtype classification of small round blue cell tumors, we demonstrate that the method is capable of finding the smallest set of genes (19 genes) with optimal classification performance, compared with results reported in the literature. CONCLUSION: In classifier design based on microarray data, the probability value derived from gene selection based on multiple combinations of data samples enables an effective mechanism for reducing the tendency of fitting local data particularities.
BACKGROUND: large scale and reliable proteins' functional annotation is a major challenge in modern biology. Phylogenetic analyses have been shown to be important for such tasks. However, up to now, phylogenetic annotation did not take into account expression data (i.e. ESTs, Microarrays, SAGE, ...). Therefore, integrating such data, like ESTs in phylogenetic annotation could be a major advance in post genomic analyses. We developed an approach enabling the combination of expression data and phylogenetic analysis. To illustrate our method, we used an example protein family, the peptidyl arginine deiminases (PADs), probably implied in Rheumatoid Arthritis. RESULTS: the analysis was performed as follows: we built a phylogeny of PAD proteins from the NCBI's NR protein database. We completed the phylogenetic reconstruction of PADs using an enlarged sequence database containing translations of ESTs contigs. We then extracted all corresponding expression data contained in EST database This analysis allowed us 1/To extend the spectrum of homologs-containing species and to improve the reconstruction of genes' evolutionary history. 2/To deduce an accurate gene expression pattern for each member of this protein family. 3/To show a correlation between paralogous sequences' evolution rate and pattern of tissular expression. CONCLUSION: coupling phylogenetic reconstruction and expression data is a promising way of analysis that could be applied to all multigenic families to investigate the relationship between molecular and transcriptional evolution and to improve functional annotation.
Several streptococcal strains had an uncharacterized mechanism of macrolide resistance that differed from those that had been reported previously in the literature. This novel mechanism conveyed resistance to 14- and 15-membered macrolides, but not to 16-membered macrolides, lincosamides or analogues of streptogramin B. The gene encoding this phenotype was cloned by standard methods from total genomic digests of Streptococcus pyogenes 02C1064 as a 4.7 kb heterologous insert into the low-copy vector, pACYC177, and expressed in several Escherichia coli K-12 strains. The location of the macrolide-resistance determinant was established by functional analysis of deletion derivatives and sequencing. A search for homologues in the genetic databases confirmed that the gene is a novel one with homology to membrane-associated pump proteins. The macrolide-resistance coding sequence was subcloned into a pET23a vector and expressed from the inducible T7 promoter on the plasmid in E. coli BL21(DE3). Physiological studies of the cloned determinant, which has been named mefA for macrolide efflux, provide evidence for its mechanism of action in host bacteria. E.coli strains containing the cloned determinant maintain lower levels of intracellular erythromycin when this compound is added to the external medium than isogenic clones without mefA. Furthermore, intracellular accumulation of [14C]-erythromycin in the original S. pyogenes strain was always lower than that observed in erythromycin-sensitive strains. This is consistent with a hypothesis that the gene encodes a novel antiporter function which pumps erythromycin out of the cell. The gene appears to be widely distributed in S. pyogenes strains, as demonstrated by primer-specific synthesis using the polymerase chain reaction.
OBJECTIVE: To determine if a false-positive trisomy 18 multiple-marker screening test (all three analytes low: maternal serum alpha-fetoprotein [AFP] at most 0.75 multiples of the median [MoM], unconjugated estriol at most 0.60 MoM, and hCG at most 0.55 MoM) indicates increased risk for obstetric complications or is related to maternal weight. METHODS: We accessed our genetic database to obtain multiple-marker screening test results, fetal karyotypes, and pregnancy outcomes from all patients with a normal multiple-marker screening test (n = 3900) and from all patients with a positive trisomy 18 screening test (n = 103) seen in the prenatal diagnosis clinic from 1992 to 1996. During this period, only maternal serum AFP was adjusted for maternal weight. RESULTS: A positive trisomy 18 screen identified five of 12 trisomy 18 fetuses. Women with a false-positive trisomy 18 screen were heavier (175.6 +/- 43.8 lb versus 159.9 +/- 37.9 lb, P < .001) and younger (29.7 +/- 6.5 years versus 32.3 +/- 6.5 years, P < .001) than women with a normal multiple-marker screening test, but were not at increased risk for pregnancy complications. Weight-adjusting all three analytes reduced the false-positive trisomy 18 screen rate by 42% (from 1.9% to 1.1%) but did not change the trisomy 18 detection rate. CONCLUSION: A false-positive trisomy 18 screening test does not indicate increased risk to develop pregnancy complications and may be related to inadequate correction for increased maternal weight.