Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

DNA-DNA hybridization values and their relationship to whole-genome sequence similarities.

DNA-DNA hybridization (DDH) values have been used by bacterial taxonomists since the 1960s to determine relatedness between strains and are still the most important criterion in the delineation of bacterial species. Since the extent of hybridization between a pair of strains is ultimately governed by their respective genomic sequences, we examined the quantitative relationship between DDH values and genome sequence-derived parameters, such as the average nucleotide identity (ANI) of common genes and the percentage of conserved DNA. A total of 124 DDH values were determined for 28 strains for which genome sequences were available. The strains belong to six important and diverse groups of bacteria for which the intra-group 16S rRNA gene sequence identity was greater than 94 %. The results revealed a close relationship between DDH values and ANI and between DNA-DNA hybridization and the percentage of conserved DNA for each pair of strains. The recommended cut-off point of 70 % DDH for species delineation corresponded to 95 % ANI and 69 % conserved DNA. When the analysis was restricted to the protein-coding portion of the genome, 70 % DDH corresponded to 85 % conserved genes for a pair of strains. These results reveal extensive gene diversity within the current concept of "species". Examination of reciprocal values indicated that the level of experimental error associated with the DDH method is too high to reveal the subtle differences in genome size among the strains sampled. It is concluded that ANI can accurately replace DDH values for strains for which genome sequences are available.

Bacterial Typing Techniques↗

ParameciumDB: a community resource that integrates the Paramecium tetraurelia genome sequence with genetic data.

ParameciumDB (http://paramecium.cgm.cnrs-gif.fr) is a new model organism database associated with the genome sequencing project of the unicellular eukaryote Paramecium tetraurelia. Built with the core components of the Generic Model Organism Database (GMOD) project, ParameciumDB currently contains the genome sequence and annotations, linked to available genetic data including the Gif Paramecium stock collection. It is thus possible to navigate between sequences and stocks via the genes and alleles. Phenotypes, of mutant strains and of knockdowns obtained by RNA interference, are captured using controlled vocabularies according to the Entity-Attribute-Value model. ParameciumDB currently supports browsing of phenotypes, alleles and stocks as well as querying of sequence features (genes, UniProt matches, InterPro domains, Gene Ontology terms) and of genetic data (phenotypes, stocks, RNA interference experiments). Forms allow submission of RNA interference data and some bioinformatics services are available. Future ParameciumDB development plans include coordination of human curation of the near 40 000 gene models by members of the research community.

Alleles↗

Protecting genomic sequence anonymity with generalization lattices.

OBJECTIVES: Current genomic privacy technologies assume the identity of genomic sequence data is protected if personal information, such as demographics, are obscured, removed, or encrypted. While demographic features can directly compromise an individual's identity, recent research demonstrates such protections are insufficient because sequence data itself is susceptible to re-identification. To counteract this problem, we introduce an algorithm for anonymizing a collection of person-specific DNA sequences. METHODS: The technique is termed DNA lattice anonymization (DNALA), and is based upon the formal privacy protection schema of k -anonymity. Under this model, it is impossible to observe or learn features that distinguish one genetic sequence from k-1 other entries in a collection. To maximize information retained in protected sequences, we incorporate a concept generalization lattice to learn the distance between two residues in a single nucleotide region. The lattice provides the most similar generalized concept for two residues (e.g. adenine and guanine are both purines). RESULTS: The method is tested and evaluated with several publicly available human population datasets ranging in size from 30 to 400 sequences. Our findings imply the anonymization schema is feasible for the protection of sequences privacy. CONCLUSIONS: The DNALA method is the first computational disclosure control technique for general DNA sequences. Given the computational nature of the method, guarantees of anonymity can be formally proven. There is room for improvement and validation, though this research provides the groundwork from which future researchers can construct genomics anonymization schemas tailored to specific datasharing scenarios.

Algorithms↗

The complete Corynebacterium glutamicum ATCC 13032 genome sequence and its impact on the production of L-aspartate-derived amino acids and vitamins.

The complete genomic sequence of Corynebacterium glutamicum ATCC 13032, well-known in industry for the production of amino acids, e.g. of L-glutamate and L-lysine was determined. The C. glutamicum genome was found to consist of a single circular chromosome comprising 3282708 base pairs. Several DNA regions of unusual composition were identified that were potentially acquired by horizontal gene transfer, e.g. a segment of DNA from C. diphtheriae and a prophage-containing region. After automated and manual annotation, 3002 protein-coding genes have been identified, and to 2489 of these, functions were assigned by homologies to known proteins. These analyses confirm the taxonomic position of C. glutamicum as related to Mycobacteria and show a broad metabolic diversity as expected for a bacterium living in the soil. As an example for biotechnological application the complete genome sequence was used to reconstruct the metabolic flow of carbon into a number of industrially important products derived from the amino acid L-aspartate.

Amino Acid Sequence↗

The genome sequence of Schoenoplectus triqueter (L.) Palla, 1888 (Poales: Cyperaceae).

We present a genome assembly of Schoenoplectus triqueter (Triangular Club-rush; Streptophyta; Magnoliopsida; Poales; Cyperaceae). The genome sequence has a total length of 580.26 megabases. Most of the assembly (98.75%) is scaffolded into 21 chromosomal pseudomolecules. Five mitochondrial sequences and the plastid genome were also assembled. Gene annotation of this assembly on Ensembl identified 31 746 protein-coding genes. This assembly was generated as part of the Darwin Tree of Life project, which produces reference genomes for eukaryotic species found in Britain and Ireland.

Poales↗

Genomic sequence and receptor for the Vibrio cholerae phage KSF-1phi: evolutionary divergence among filamentous vibriophages mediating lateral gene transfer.

KSF-1phi, a novel filamentous phage of Vibrio cholerae, supports morphogenesis of the RS1 satellite phage by heterologous DNA packaging and facilitates horizontal gene transfer. We analyzed the genomic sequence, morphology, and receptor for KSF-1phi infection, as well as its phylogenetic relationships with other filamentous vibriophages. While strains carrying the mshA gene encoding mannose-sensitive hemagglutinin (MSHA) type IV pilus were susceptible to KSF-1phi infection, naturally occurring MSHA-negative strains and an mshA deletion mutant were resistant. Furthermore, d-mannose as well as a monoclonal antibody against MSHA inhibited infection of MSHA-positive strains by the phage, suggesting that MSHA is the receptor for KSF-1phi. The phage genome comprises 7,107 nucleotides, containing 14 open reading frames, 4 of which have predicted protein products homologous to those of other filamentous phages. Although the overall genetic organization of filamentous phages appears to be preserved in KSF-1phi, the genomic sequence of the phage does not have a high level of identity with that of other filamentous phages and reveals a highly mosaic structure. Separate phylogenetic analysis of genomic sequences encoding putative replication proteins, receptor-binding proteins, and Zot-like proteins of 10 different filamentous vibriophages showed different results, suggesting that the evolution of these phages involved extensive horizontal exchange of genetic material. Filamentous phages which use type IV pili as receptors were found to belong to different branches. While one of these branches is represented by CTXphi, which uses the toxin-coregulated pilus as its receptor, at least four evolutionarily diverged phages share a common receptor MSHA, and most of these phages mediate horizontal gene transfer. Since MSHA is present in a wide variety of V. cholerae strains and is presumed to express in the environment, diverse filamentous phages using this receptor are likely to contribute significantly to V. cholerae evolution.

Bacteriophages↗

A microcosting and cost consequence analysis from a randomized controlled trial comparing genome sequencing with exome sequencing for genetic diagnosis.

PURPOSE: Diagnosing rare diseases is costly. The objectives were to microcost exome (ES) and genome sequencing (GS) trios and estimate the incremental costs of GS per additional diagnosis from an institutional payer perspective. METHODS: Trios (proband plus biological parents) that are referred for sequencing were randomly assigned to ES or GS. Laboratory workflow and sequencing were microcosted. Total and category cost per trio were estimated probabilistically. Effectiveness was expressed as diagnostic yield (rates of diagnostic or partially diagnostic variants detected). Incremental costs and effectiveness were calculated. RESULTS: The mean total cost per trio was CAD 2888.79 (95% CI 2567.72, 3492.72) for ES (n = 329) and 4364.02 (95% CI 3984.94, 5013.67) for GS (n = 324). Reagents accounted for 34% and 61% of total costs for ES and GS, respectively. The incremental cost of GS was 1475.23. The diagnostic yield was 35.9% for ES and 32.7% for GS with a difference of 0.032 (95% CI: -0.041, 0.104, P value .397). CONCLUSION: GS demonstrated higher costs and a similar diagnostic yield to ES but was limited by technical capabilities at the time of the study. The study provides comprehensive costs for the economic evaluation comparing alternative diagnostic pathways and impetus for further evaluating variants uniquely detectable by GS.

Humans↗

Metallothionein cDNA, promoter, and genomic sequences of the tropical green mussel, Perna viridis.

The primary structure of the cDNA and metallothionein (MT) genomic sequences of the tropical green mussel (Perna viridis) was determined. The complete cDNA sequences were obtained using degenerate primers designed from known metallothionein consensus amino acid sequences from the temperate species Mytilus edulis. The amino acid sequences of P. viridis metallothionein deduced from the coding region consisted of 72 amino acids with 21 cysteine residues and 9 Cys-X-Cys motifs corresponding to Type I MT class of other species. Two different genomic sequences coding for the same mRNA were obtained. Each putative gene contained a unique 5'UTR and two unique introns located at the same splice sites. The promoters for both genes were different in length and both contained metal responsive elements and active protein-binding sites. The structures of the genomic clones were compared with those of other species. J. Exp. Zool. 284:445-453, 1999.

Amino Acid Sequence↗

Using the chicken genome sequence in the development and mapping of genetic markers in the turkey (Meleagris gallopavo).

The efficacy of employing the chicken genome sequence in developing genetic markers and in mapping the turkey genome was studied. Eighty previously uncharacterized microsatellite markers were identified for the turkey using BLAST alignment to the chicken genome. The chicken sequence was then used to develop primers for polymerase chain reaction where the turkey sequence was either unavailable or insufficient. A total of 78 primer sets were tested for amplification and polymorphism in the turkey, and informative markers were genetically mapped. Sixty-five (83%) amplified turkey genomic DNA, and 33 (42%) were polymorphic in the University of Minnesota/Nicholas Turkey Breeding Farms mapping families. All but one marker genetically mapped to the position predicted from the chicken genome sequence. These results demonstrate the usefulness of the chicken sequence for the development of genomic resources in other avian species.

Alleles↗

Porcine Fas-ligand gene: genomic sequence analysis and comparison with human gene.

Thymic Fas-ligand (FasL) cDNA and hepatic FasL genomic sequences were obtained from a 2-month-old LW pig. From these nucleotide sequences, amino acid sequence was deduced and compared with FasL sequences obtained from various animals. This comparison reveals that porcine FasL is closer to that of human, macaca and cat, and differs more from mouse and rat. The extracelluar domains of porcine and human FasL proteins appear to be functionally compatible. The complete genomic DNA sequence of porcine FasL was also compared with its human counterpart. Exons showed 80-89% nucleotide homology between pig and human, while introns showed 64-69% nucleotide homology. Sequence comparison by Harr plot analysis revealed many stretches within introns having identical sequences, suggesting that the sites may have unidentified common functions. One potential extra exon between exons 2 and 3 was located within porcine intron 2. This potential exon has no counterpart in human FasL intron 2. Whether or not this extra exon can be expressed and could cause additional immunological responses remains to be investigated. For future xenotransplantation, it is important to compare porcine and human genomic sequences, and to investigate their system compatibilities.

Amino Acid Sequence↗

Genome sequence of Avery's virulent serotype 2 strain D39 of Streptococcus pneumoniae and comparison with that of unencapsulated laboratory strain R6.

Streptococcus pneumoniae (pneumococcus) is a leading human respiratory pathogen that causes a variety of serious mucosal and invasive diseases. D39 is an historically important serotype 2 strain that was used in experiments by Avery and coworkers to demonstrate that DNA is the genetic material. Although isolated nearly a century ago, D39 remains extremely virulent in murine infection models and is perhaps the strain used most frequently in current studies of pneumococcal pathogenesis. To date, the complete genome sequences have been reported for only two S. pneumoniae strains: TIGR4, a recent serotype 4 clinical isolate, and laboratory strain R6, an avirulent, unencapsulated derivative of strain D39. We report here the genome sequences and new annotation of two different isolates of strain D39 and the corrected sequence of strain R6. Comparisons of these three related sequences allowed deduction of the likely sequence of the D39 progenitor and mutations that arose in each isolate. Despite its numerous repeated sequences and IS elements, the serotype 2 genome has remained remarkably stable during cultivation, and one of the D39 isolates contains only five relatively minor mutations compared to the deduced D39 progenitor. In contrast, laboratory strain R6 contains 71 single-base-pair changes, six deletions, and four insertions and has lost the cryptic pDP1 plasmid compared to the D39 progenitor strain. Many of these mutations are in or affect the expression of genes that play important roles in regulation, metabolism, and virulence. The nature of the mutations that arose spontaneously in these three strains, the relative global transcription patterns determined by microarray analyses, and the implications of the D39 genome sequences to studies of pneumococcal physiology and pathogenesis are presented and discussed.

Animals↗

Complete genome sequence of Streptomyces californicus ADR1, an anti-infective, anti-biofilm and anti-oxidant producing endophyte isolated from the medicinal plant Datura metel.

OBJECTIVE: Streptomyces californicus strain ADR1 is an endophytic actinobacterium isolated from Datura metel that produces secondary metabolites with potent antibacterial and anti-biofilm activities against WHO-listed high-priority Gram-positive pathogens. While anti-bacterial and antioxidant potential of the strain ADR1 has been extensively characterized, its complete genome sequence remains to be investigated for further insights into its biosynthetic potential. This study presents the complete genome sequence analysis of the strain ADR1 to provide a robust genomic foundation for understanding its metabolic versatility and biosynthesis of compounds with therapeutic significance. DATA DESCRIPTION: The ADR1 genome was sequenced using Illumina HiSeq. The assembly comprised 262 scaffolds with a total genome size of 8.4 Mb and G + C content of 72.5%, containing 7427 protein-coding genes. AntiSMASH and IIT-Hyderabad novelBGC analysis revealed 39 biosynthetic gene clusters, including non-ribosomal peptide synthetases, type I polyketide synthases, terpene and melanin clusters, correlating with the diverse therapeutic compounds previously identified through GC-MS analysis. This high-quality genome provides crucial insights into the biosynthetic potential underlying potent antimicrobial and antioxidant activities of the strain ADR1.

Streptomyces↗

The Drosophila melanogaster 60A chromosomal division is extremely dense with functional genes: their sequences, genomic organization, and expression.

We cloned and sequenced genomic DNA contigs spanning over 45 kb, surrounding the insertion site of the P-element that is responsible for the developmental defects in the ken and barbie (ken) mutant of Drosophila. This region harbors 10 functional transcription units, in addition to the already well-characterized TGFbeta-60A gene. These include the genes, undefined 1 (UD1), UD2, and UD3, each coding for proteins of unknown function, the ken gene encoding a new Krüppel-like putative transcription factor, the fly homologues of the mammalian mitochondrial trifunctional enzyme (thiolase), and the TAR DNA-binding protein-43 (TBPH), the first nonvertebrate member of the transmembrane 4 superfamily (TM4SF) gene, a new homeodomain gene, and a gene coding for a putative nuclear binding protein (PNBP) that is homologous to maleless, and a Copia-like element. UD3 exists in an intron of the maleless homologue, yet is expressed independent of it. The UD1 and TM4SF genes orient in a tail-to-tail manner with their 3' untranslated region sequences overlapping over 44 nucleotides. Thus the partial overlap and intraintronic organization permitted dense packing of the functional genes within a short segment of the genome.

Acetyl-CoA C-Acetyltransferase↗

A computer simulation analysis of the accuracy of partial genome sequencing and restriction fragment analysis in estimating genetic relationships: an application to papillomavirus DNA sequences.

BACKGROUND: Determination of genetic relatedness among microorganisms provides information necessary for making inferences regarding phylogeny. However, there is little information available on how well the genetic relationships inferred from different genotyping methods agree with true genetic relationships. In this report, two genotyping methods - restriction fragment analysis (RFA) and partial genome DNA sequencing - were each compared to complete DNA sequencing as the definitive standard for classification. RESULTS: Using the Genbank database, 16 different types or subtypes of papillomavirus were selected as study samples, because numerous complete genome sequences were available. RFA was achieved by computer-simulated digestion. The genetic similarity of samples, based on RFA, was determined from the proportion of fragments that matched in size. DNA sequences of four specific genes (E1, E6, E7, and L1), representing partial genome sequencing, were also selected for comparison to complete genome sequencing. Laboratory error was not taken into account. Evaluation of the correlation between genetic similarity matrices (Mantel's r) and comparisons of the structure of the derived dendrograms (partition metric) indicated that partial genome sequencing (for single genes) had higher agreement with complete genome sequencing, achieving a maximum Mantel's r = 0.97 and a minimum partition metric = 10. RFA had lower agreement, with a maximum Mantel's r = 0.60 and a minimum partition metric = 18. CONCLUSIONS: This simulation indicated that for smaller genomes, such as papillomavirus, partial genome sequencing is superior to restriction fragment analysis in representing genetic relatedness among isolates. The generalizability of these results to larger genomes, as well as the impact of laboratory error, remains to be demonstrated.

Animals↗

Bisulfite genomic sequencing: systematic investigation of critical experimental parameters.

Bisulfite genomic sequencing is the method of choice for the generation of methylation maps with single-base resolution. The method is based on the selective deamination of cytosine to uracil by treatment with bisulfite and the sequencing of subsequently generated PCR products. In contrast to cytosine, 5-methylcytosine does not react with bisulfite and can therefore be distinguished. In order to investigate the potential for optimization of the method and to determine the critical experimental parameters, we determined the influence of incubation time and incubation temperature on the deamination efficiency and measured the degree of DNA degradation during the bisulfite treatment. We found that maximum conversion rates of cytosine occurred at 55 degrees C (4-18 h) and 95 degrees C (1 h). Under these conditions at least 84-96% of the DNA is degraded. To study the impact of primer selection, homologous DNA templates were constructed possessing cytosine-containing and cytosine-free primer binding sites, respectively. The recognition rates for cytosine (>/=97%) and 5-methylcytosine (>/=94%) were found to be identical for both templates.

5-Methylcytosine↗

Third-generation whole-genome sequencing reveals the role of CNTNAP2 as a tumor suppressor gene in high-risk neuroblastomas.

BACKGROUND: Neuroblastoma is a common and aggressive pediatric sympathetic nervous system tumor. Genomic structural variants (SVs) contribute substantially to neuroblastoma, yet remain under-characterized in high-risk neuroblastomas. We aimed to elucidate neuroblastoma pathogenesis using third-generation whole-genome sequence high-risk cases to identify driver aberrations and explore potential therapeutic strategies. METHODS: We analyzed third-generation whole-genome sequencing data of 20 high-risk neuroblastoma samples and combined the findings with those obtained from the analysis of clinical samples, in vitro models, and public datasets. RESULTS: The contactin-associated protein-like 2 (CNTNAP2) gene was observed to be frequently aberrated because of structural variants in high-risk neuroblastoma samples. CNTNAP2 expression was significantly correlated with favorable histology and could be used to predict prognosis using clinical samples and neuroblastoma datasets. Overexpression and knockdown experiments and transcriptomic analysis revealed that CNTNAP2 was primarily involved in neuronal differentiation and axon guidance pathways; moreover, CNTNAP2 was required for neuroblastoma differentiation and affected cancer stemness. Immunoprecipitation and mass spectrometry revealed that CNTNAP2 interacted with cytoskeletal proteins like drebrin 1 (DBN1) and myosin-heavy chain 9 (MYH9). CNTNAP2 dynamically reorganises actin and microtubules for DBN1-mediated neuronal differentiation. CNTNAP2 also reduces CTNNB1 transcription and β-catenin pathway activation by inhibiting MYH9 nuclear translocation. CNTNAP2 overexpression in neuroblastoma cell lines resulted in cell cycle arrest, decreased cell proliferation and metastasis. CONCLUSIONS: The recurrent loss of CNTNAP2 in neuroblastoma contributes to an aggressive phenotype by impairing neuronal differentiation and increasing cancer stemness. These findings may serve as a foundation for developing therapeutic strategies to overcome barriers to differentiation.

Humans↗

Genomic sequences encoding the acidic and basic subunits of Mojave toxin: unusually high sequence identity of non-coding regions.

Mojave toxin (Mtx) is a heterodimeric, neurotoxic phospholipase A2 (PLA2) found in the venom of the Mojave rattlesnake, Crotalus scutulatus scutulatus, and is characteristic of all rattlesnake presynaptic neurotoxins. This paper describes the isolation and nucleotide (nt) sequence of the genomic clones encoding both the non-neurotoxic, non-enzymatic acidic subunit (Mtx-a) and the toxic, PLA2-active basic subunit (Mtx-b), and compares their structures. Both cloned genes shared virtually identical overall organization, with four exons separated by three introns, which were inserted in the same relative positions of the genes' coding regions. The exon/intron structure was similar to that reported for mammalian PLA2 genes. Most remarkable was the high degree of nt sequence identity between Mtx-a and Mtx-b. While the exons shared about 70% identity, the introns were greater than 90% identical and the 5' and 3' untranslated and flanking regions were greater than 95% identical. These findings support our earlier suggestion [Aird et al., Biochemistry 24 (1985) 7054-7058] that the genes coding for the two subunits arose from a common ancestor. There has clearly been a strong selection on the nt sequence of the non-coding regions during this evolutionary process. This is the first report of genomic sequences of PLA2-like proteins from snakes.

Amino Acid Sequence↗