Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Delivering effective genome sequencing in pediatric care: From research in the 100,000 Genomes Project to routine clinical practice.

PURPOSE: Genome sequencing (GS) is increasingly used to investigate rare conditions, primarily in children. The 100,000 Genomes Project (100KG) evaluated GS ahead of implementation in the English National Health Service. In 2020, the National Health Service Genomic Medicine Service (GMS) became the first public health care system to offer GS in routine clinical care. We investigate how learning from 100KG informed GMS service delivery. METHODS: We compare GS outcomes in children tested at a large pediatric hospital via GMS (n = 501) and 100KG research (n = 1759). RESULTS: GMS diagnostic yield (29%) was higher than that in 100KG (22%) (P < .0016). Median age at testing was 8 years in 100KG and 6 in the GMS (P < .05). In 100KG, the diagnostic yield was <10% for 15 indications, none of which are included in GMS testing. 100KG data showed little benefit to application of >3 panels. Use of fewer but larger GMS panels resulted in a significantly higher number of genes tested per patient: median 2801 vs 1373 in 100KG (P < .001). In 100KG, diagnostic yield was not significantly increased by testing more than 3 family members (n = 34/142, 24%). CONCLUSION: Learning from 100KG has informed GS clinical service delivery, resulting in higher diagnostic yields and earlier age at testing. Lessons are broadly applicable to all services providing GS, enabling earlier access to tailored management with fewer investigations.

Humans↗

Distribution and intensity of constraint in mammalian genomic sequence.

Comparisons of orthologous genomic DNA sequences can be used to characterize regions that have been subject to purifying selection and are enriched for functional elements. We here present the results of such an analysis on an alignment of sequences from 29 mammalian species. The alignment captures approximately 3.9 neutral substitutions per site and spans approximately 1.9 Mbp of the human genome. We identify constrained elements from 3 bp to over 1 kbp in length, covering approximately 5.5% of the human locus. Our estimate for the total amount of nonexonic constraint experienced by this locus is roughly twice that for exonic constraint. Constrained elements tend to cluster, and we identify large constrained regions that correspond well with known functional elements. While constraint density inversely correlates with mobile element density, we also show the presence of unambiguously constrained elements overlapping mammalian ancestral repeats. In addition, we describe a number of elements in this region that have undergone intense purifying selection throughout mammalian evolution, and we show that these important elements are more numerous than previously thought. These results were obtained with Genomic Evolutionary Rate Profiling (GERP), a statistically rigorous and biologically transparent framework for constrained element identification. GERP identifies regions at high resolution that exhibit nucleotide substitution deficits, and measures these deficits as "rejected substitutions". Rejected substitutions reflect the intensity of past purifying selection and are used to rank and characterize constrained elements. We anticipate that GERP and the types of analyses it facilitates will provide further insights and improved annotation for the human genome as mammalian genome sequence data become richer.

Animals↗

Proteomics identify disease-associated variants in patients with rare diseases undiagnosed after genome sequencing.

Despite the introduction of genome sequencing (GS) for rare disease diagnostics, a genetic cause is not identified in most patients. Here, we explored the potential of proteomics to improve the diagnostic yield in 424 patients with rare diseases from the 100,000 Genomes Project (100kGP) without a genetic diagnosis. Serum proteomic profiling was performed using the Olink Explore 1536 assay (N&#xa0;=&#xa0;1463 proteins). For 13 patients without genetic diagnoses, detection of lower serum protein "outliers" (z-score&#xa0;<&#xa0;-2) led to confirmed genetic diagnoses by resolving variants of uncertain significance or prioritizing genes for targeted GS reanalysis. For 23 additional patients without genetic diagnoses (64% of findings), we identified candidate gene-disease links and variants through convergent evidence from lower protein outliers and variants ranked through the variant prioritization tool Exomiser. For example, we identified a candidate heterozygous missense variant [Genome Aggregation Database (gnomAD) minor allele frequency&#xa0;=&#xa0;0.006%] in tyrosine kinase with immunoglobulin-like and epidermal growth factor homology domains 1 (TIE1) that was only present in a patient with lower TIE1 serum abundance (z-score&#xa0;=&#xa0;-5.12) and their father, both of whom were affected by the same monogenic cardiac disorder, but in no other individuals from the 100kGP. Missense (52.5%) and splice region (27.5%) variants accounted for most diagnostic or candidate variants prioritized. This proof-of-principle study demonstrated that serum proteomics can support rare disease diagnosis and identify disease-causing genes in patients undiagnosed after GS, although successful implementation will likely depend on tissue specificity of protein expression, detectability in blood, proteomic platform coverage, and sensitivity.

Humans↗

Inferences from whole-genome sequences of bacterial pathogens.

Genomic sequencing of bacterial pathogens has recently moved from the study of distantly related organisms to within-species comparisons of multiple strains. Strains often differ in their ability to cause disease, and comparative genomics is uncovering novel virulence determinants, hidden aspects of pathogenesis, and new targets for vaccine development. DNA microarrays and other gene-survey techniques are being used to quantify variability in gene content within bacterial populations, and to reveal the strain-specific basis for diversity and severity of pathology.

Bacteria↗

Genome sequence and characterization of a virus (HaRNAV) related to picorna-like viruses that infects the marine toxic bloom-forming alga Heterosigma akashiwo.

Heterosigma akashiwo (Rhaphidophyceae) is a unicellular, flagellated, bloom-forming, toxic alga of ecological and economic importance. Here, we report the results of sequencing and analyzing the genome of an 8.6-kb single-stranded RNA virus (HaRNAV-SOG263) that infects H. akashiwo. Our results show that HaRNAV is related to picorna-like viruses, but does not belong within any currently defined virus family. This is based on the genome organization and sequence comparisons of putative RNA-dependent RNA polymerase (RdRp), helicase, and capsid protein sequences. The genome sequence predicts a single open reading frame (orf) encoding a polyprotein that contains conserved picorna-like protein domains, with putative nonstructural protein domains present in the N-terminus and the structural proteins in the C-terminus of the polyprotein. We have analyzed and compared the virus structural proteins from infectious and noninfectious particles. In this way, we identified structural protein cleavage sites as well as protein processing events that are presumably important for maturation of virus particles. The combination of genome structure and sequence relationships to other viruses suggests that HaRNAV is the first member of a proposed new virus family (Marnaviridae), related to picorna-like viruses.

Amino Acid Sequence↗

Computational comparison of human genomic sequence assemblies for a region of chromosome 4.

Much of the available human genomic sequence data exist in a fragmentary draft state following the completion of the initial high-volume sequencing performed by the International Human Genome Sequencing Consortium (IHGSC) and Celera Genomics (CG). We compared six draft genome assemblies over a region of chromosome 4p (D4S394-D4S403), two consecutive releases by the IHGSC at University of California, Santa Cruz (UCSC), two consecutive releases from the National Centre for Biotechnology Information (NCBI), the public release from CG, and a hybrid assembly we have produced using IHGSC and CG sequence data. This region presents particular problems for genomic sequence assembly algorithms as it contains a large tandem repeat and is sparsely covered by draft sequences. The six assemblies differed both in terms of their relative coverage of sequence data from the region and in their estimated rates of misassembly. The CG assembly method attained the lowest level of misassembly, whereas NCBI and UCSC assemblies had the highest levels of coverage. All assemblies examined included <60% of the publicly available sequence from the region. At least 6% of the sequence data within the CG assembly for the D4S394-D4S403 region was not present in publicly available sequence data. We also show that even in a problematic region, existing software tools can be used with high-quality mapping data to produce genomic sequence contigs with a low rate of rearrangements.

Chromosomes, Human, Pair 4↗

Complete chloroplast genome sequence of Gycine max and comparative analyses with other legume genomes.

Lack of complete chloroplast genome sequences is still one of the major limitations to extending chloroplast genetic engineering technology to useful crops. Therefore, we sequenced the soybean chloroplast genome and compared it to the other completely sequenced legumes, Lotus and Medicago. The chloroplast genome of Glycine is 152,218 basepairs (bp) in length, including a pair of inverted repeats of 25,574 bp of identical sequence separated by a small single copy region of 17,895 bp and a large single copy region of 83,175 bp. The genome contains 111 unique genes, and 19 of these are duplicated in the inverted repeat (IR). Comparisons of Glycine, Lotus and Medicago confirm the organization of legume chloroplast genomes based on previous studies. Gene content of the three legumes is nearly identical. The rpl22 gene is missing from all three legumes, and Medicago is missing rps16 and one copy of the IR. Gene order in Glycine, Lotus, and Medicago differs from the usual gene order for angiosperm chloroplast genomes by the presence of a single, large inversion of 51 kilobases (kb). Detailed analyses of repeated sequences indicate that many of the Glycine repeats that are located in the intergenic spacer regions and introns occur in the same location in the other legumes and in Arabidopsis, suggesting that they may play some functional role. The presence of small repeats of psbA and rbcL in legumes that have lost one copy of the IR indicate that this loss has only occurred once during the evolutionary history of legumes.

Base Sequence↗

Genome sequencing: the ripping yarn of the frozen genome.

The completion of the genome sequence of the filamentous fungus Neurospora crassa reveals a gene number very much higher than those of yeasts. Of particular interest in this species are the effects of the repeat-induced point mutation (RIP) process, which appears to have prevented recent evolution through gene duplication in this lineage.

Evolution, Molecular↗

Methods for obtaining and analyzing whole chloroplast genome sequences.

During the past decade, there has been a rapid increase in our understanding of plastid genome organization and evolution due to the availability of many new completely sequenced genomes. There are 45 complete genomes published and ongoing projects are likely to increase this sampling to nearly 200 genomes during the next 5 years. Several groups of researchers including ours have been developing new techniques for gathering and analyzing entire plastid genome sequences and details of these developments are summarized in this chapter. The most important developments that enhance our ability to generate whole chloroplast genome sequences involve the generation of pure fractions of chloroplast genomes by whole genome amplification using rolling circle amplification, cloning genomes into Fosmid or bacterial artificial chromosome (BAC) vectors, and the development of an organellar annotation program (Dual Organellar GenoMe Annotator [DOGMA]). In addition to providing details of these methods, we provide an overview of methods for analyzing complete plastid genome sequences for repeats and gene content, as well as approaches for using gene order and sequence data for phylogeny reconstruction. This explosive increase in the number of sequenced plastid genomes and improved computational tools will provide many insights into the evolution of these genomes and much new data for assessing relationships at deep nodes in plants and other photosynthetic organisms.

Amino Acid Sequence↗

Helicobacter pylori unmasked--the complete genome sequence.

The publication of the complete genome sequence of Helicobacter pylori and the computer analysis of the genes it contains represents an enormous advance in H. pylori research. Besides providing an overview of H. pylori biology, the genome sequence will aid future research and radically alter the research process. In particular, it will enable a 'top down' approach whereby genes are investigated because of their similarity to genes of known function in other organisms. Other advances in biotechnology will allow all H. pylori genes to be studied simultaneously, proteins to be identified quickly, and potential drug targets to be evaluated efficiently. Progress in H. pylori research should now be rapid, and with the research interest and resources focused on it, H. pylori could become the paradigm for post-genomic research.

Genetics, Microbial↗

Rnai: a new technology in the post-genomic sequencing era.

The complete genome sequences and huge numbers of predicted gene sequences from many complex organisms are available. Reverse-genetic analyses of these organisms will now be necessary to understand what all these genes are doing and how their functions interact. RNAi technology is very useful in this regard for studying gene function, not only in these model organisms, but also in the organisms previously considered not to be amenable to genetic analysis. This treatment reviews the discovery of RNAi, advances in the study of the mechanisms by which this material impinges on gene-product expressions, and technical improvements that will allow RNAi to be applied to a wide variety of organisms.

Animals↗

Construction and utility of 10-kb libraries for efficient clone-gap closure for rice genome sequencing.

Rice is an important crop and a model system for monocot genomics, and is a target for whole genome sequencing by the International Rice Genome Sequencing Project (IRGSP). The IRGSP is using a clone by clone approach to sequence rice based on minimum tiles of BAC or PAC clones. For chromosomes 10 and 3 we are using an integrated physical map based on two fingerprinted and end-sequenced BAC libraries to identifying a minimum tiling path of clones. In this study we constructed and tested two rice genomic libraries with an average insert size of 10 kb (10-kb library) to support the gap closure and finishing phases of the rice genome sequencing project. The HaeIII library contains 166,752 clones covering approximately 4.6x rice genome equivalents with an average insert size of 10.5 kb. The Sau3AI library contains 138,960 clones covering 4.2x genome equivalents with an average insert size of 11.6 kb. Both libraries were gridded in duplicate onto 11 high-density filters in a 5 x 5 pattern to facilitate screening by hybridization. The libraries contain an unbiased coverage of the rice genome with less than 5% contamination by clones containing organelle DNA or no insert. An efficient method was developed, consisting of pooled overgo hybridization, the selection of 10-kb gap spanning clones using end sequences, transposon sequencing and utilization of in silico draft sequence, to close relatively small gaps between sequenced BAC clones. Using this method we were able to close a majority of the gaps (up to approximately 50 kb) identified during the finishing phase of chromosome-10 sequencing. This method represents a useful way to close clone gaps and thus to complete the entire rice genome.

Chromosomes, Artificial, Bacterial↗

Whole genome sequencing and phylogenetic classification accelerate the implementation of respiratory syncytial virus genomic surveillance in Canada: a pilot study.

UNLABELLED: Whole genome sequencing (WGS) has emerged as a powerful tool to facilitate the study of existing and emerging infectious diseases. WGS-based genomic surveillance provides information on the genetic diversity and tracks the evolution of important viral pathogens, including respiratory syncytial virus (RSV). Multiplex tiling polymerase chain reaction (PCR) assays have been used to facilitate sequencing of a variety of pathogens in support of genomics-based surveillance initiatives. We developed, optimized, and implemented multiplex tiling PCR assays for RSVA and RSVB capable of generating near-complete genomes in the majority of contemporaneous specimens tested. A pilot data set comprising 52 RSVA and 37 RSVB genomes derived from Canadian clinical specimens during the 2022-2023 respiratory virus season was used to perform phylogenetic analyses using both near-complete genome and glycoprotein (G) sequences. Overall, the RSV phylogenetic tree built with whole genomes showed identical lineage clusters as compared to the G gene but was more discriminatory. Moreover, the availability of complete genomes enables the identification of a broader range of mutations. For instance, mutations identified in the fusion protein among Canadian isolates tested here, including S377N, K272M, S276N, S211N, S206I, and S209Q, could affect the efficacy of current vaccines or antiviral-based therapeutics. In conclusion, our work reinforces other recent studies demonstrating the utility of multiplex tiling PCR assays to facilitate high-throughput WGS of RSV, which is capable of supporting enhanced genomic surveillance initiatives, as well as the more comprehensive genomic analyses required to inform public health strategies for the development and usage of vaccines and antiviral drugs. IMPORTANCE: We present assays to efficiently sequence genomes of RSVA and RSVB. This enables researchers and public health agencies to acquire high-quality genomic data using rapid and cost-effective approaches. Genomic data-based comparative analysis can be used to conduct surveillance and monitor circulating isolates for efficacy of vaccines and antiviral therapeutics.

Humans↗

Molecular residual disease assessment in colorectal and bladder cancer by somatic structural variant analysis of cell-free DNA whole-genome sequencing data.

BACKGROUND: Whole-genome sequencing (WGS)-based methods for circulating tumor DNA (ctDNA) detection typically rely on tumor-informed identification of somatic single nucleotide variants (SNVs). Somatic structural variants (SVs) are another type of cancer-specific genomic alteration, which owing to their larger genomic footprint and unique breakpoint junctions, are easier to distinguish from sequencing noise than SNVs. They are, however, rarely used for ctDNA detection because of (1) artifacts from WGS procedures that SV callers may falsely interpret as genuine SVs. This makes it difficult to establish high-confidence SV catalogos from short-read tumor WGS and can cause false-positive ctDNA detections. (2) Lack of robust strategies to quantify SV-supporting reads in plasma WGS. To address these barriers and enable integration of SV biomarkers into WGS-based ctDNA detection, we present a bioinformatic framework for algorithmic curation of somatic SV calls from fresh-frozen and formalin-fixed paraffin-embedded (FFPE) tumors, coupled with a novel approach for sensitive, accurate mapping and quantification of SV breakpoint-supporting reads in plasma WGS. METHODS: Tumor, normal and plasma WGS data from 144 patients with stage III colorectal cancer was used to establish the bioinformatic framework. This included ~30x WGS data from 1564 serially collected plasma samples. The framework was validated using tumor/normal/plasma WGS data from 32 patients with muscle-invasive bladder cancer. SV-based ctDNA detection was benchmarked against previously published SNV-based ctDNA results for the same samples. RESULTS: After curation of SV calls and quantification in plasma WGS, our SV-based approach enabled robust ctDNA detection with overall specificity exceeding 99% in plasma samples. Furthermore, we observed strong concordance (Pearson&#x2019;s r&#x2009;>&#x2009;0.93, p&#x2009;<&#x2009;2.2&#x2009;&#xd7;&#x2009;10&#x2212; 16) between ctDNA-positive samples identified by our SV-based method and previous SNV-based analyses, validating the reliability of our approach. Finally, we demonstrated application of the method in an independent bladder cancer cohort, highlighting its generalizability and potential clinical use. CONCLUSIONS: We provide a bioinformatic framework that establishes somatic SVs as ultra-specific biomarkers for WGS-based, tumor-informed ctDNA detection. The approach delivers specific detection even when the SV catalogos are established from FFPE samples. The SV framework can stand alone or enhance SNV-based analysis pipelines.

Humans↗

High-speed conversion of cytosine to uracil in bisulfite genomic sequencing analysis of DNA methylation.

Bisulfite genomic sequencing is a widely used technique for analyzing cytosine-methylation of DNA. By treating DNA with bisulfite, cytosine residues are deaminated to uracil, while leaving 5-methylcytosine largely intact. Subsequent PCR and nucleotide sequence analysis permit unequivocal determination of the methylation status at cytosine residues. A major caveat associated with the currently practiced procedure is that it takes 16-20 hr for completion of the conversion of cytosine to uracil. Here we report that a complete deamination of cytosine to uracil can be achieved in shorter periods by using a highly concentrated bisulfite solution at an elevated temperature. Time course experiments demonstrated that treating DNA with 9 M bisulfite for 20 min at 90 degrees C or 40 min at 70 degrees C all cytosine residues in the DNA were converted to uracil. Under these conditions, the majority of 5-methylcytosines remained intact. When a high molecular weight DNA derived from a cell line (containing a number of genes whose methylation status was known) was treated with bisulfite under the above conditions and amplified and sequenced, the results obtained were consistent with those reported in the literature. Although some degradation of DNA occurred during this process, the amount of treated DNA required for the amplification was nearly equal to that required for the conventional bisulfite genomic sequencing procedure. The increased speed of DNA methylation analysis with this novel procedure is expected to advance various aspects of DNA sciences.

Base Sequence↗

Yeast genome sequencing: the power of comparative genomics.

For decades, unicellular yeasts have been general models to help understand the eukaryotic cell and also our own biology. Recently, over a dozen yeast genomes have been sequenced, providing the basis to resolve several complex biological questions. Analysis of the novel sequence data has shown that the minimum number of genes from each species that need to be compared to produce a reliable phylogeny is about 20. Yeast has also become an attractive model to study speciation in eukaryotes, especially to understand molecular mechanisms behind the establishment of reproductive isolation. Comparison of closely related species helps in gene annotation and to answer how many genes there really are within the genomes. Analysis of non-coding regions among closely related species has provided an example of how to determine novel gene regulatory sequences, which were previously difficult to analyse because they are short and degenerate and occupy different positions. Comparative genomics helps to understand the origin of yeasts and points out crucial molecular events in yeast evolutionary history, such as whole-genome duplication and horizontal gene transfer(s). In addition, the accumulating sequence data provide the background to use more yeast species in model studies, to combat pathogens and for efficient manipulation of industrial strains.

Biological Evolution↗

Rapid full-length genomic sequencing of two cytopathically heterogeneous Australian primary HIV-1 isolates.

Two Australian HIV-1 isolates, derived from patient blood (HIV(MBC200)) and cerebrospinal fluid (HIV(MBC925)), were characterized after in vitro culture in peripheral blood mononuclear cells (PBMC). Although virus replication was similar, as measured by cell-free reverse transcriptase activity, only one of the two isolates (HIV-1(MCB200)) consistently induced cell syncytia and depleted the PBMC population of CD4+ cells by cell killing. A novel technique, devised for rapidly obtaining high-quality viral sequence data and the full-length genomic sequence of these two isolates, is presented. Analysis of the predicted sequence of the viral Env proteins provides correlates of the observed phenotypes. Phylogenetic analysis derived using near full-length sequence of these and other HIV-1 subtype B genomic sequences (including two other Australian isolates) shows a star-shaped phylogeny with each member having a similar genetic diversity. These data expand the database of genomic sequence available from well-characterized primary clinical isolates of HIV-1 using a novel rapid technique.

Acquired Immunodeficiency Syndrome↗

Genomic sequencing by ligation mediated polymerase chain reaction using direct blotting and non-radioactive detection.

Genomic sequencing has become an important tool for analyzing uncloned cellular DNA with regard to the methylation status of cytidines as well as to DNA-protein interactions within cells. The hybridization step of the genomic sequencing procedure requires a very high sensitivity, rendering the method fairly difficult. Using a modified ligation mediated polymerase chain reaction procedure (LMPCR) and a sensitive non-radioactive detection method, we have developed a procedure avoiding the high amounts of radioactivity formerly needed for detection of chemically cleaved genomic DNA. The detection limit of our method of genomic sequencing is less than 1 microgram mammalian DNA, which is much better than the detection limit of the original genomic sequencing method and comparable with the detection limit of radioactive detection after the LMPCR procedure. In addition to the advantages of the non-radioactive detection technique we simplified the blotting step of the genomic sequencing procedure by using the direct blotting electrophoresis method. The method was applied to a region 5' to the human c-myc promoter of HeLa cells and was able to verify the sequence obtained by other authors and to specify the methylation status of five CpG-pairs within this sequence.

Base Sequence↗