Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Evaluation of annotation strategies using an entire genome sequence.

MOTIVATION: Genome-wide functional annotation either by manual or automatic means has raised considerable concerns regarding the accuracy of assignments and the reproducibility of methodologies. In addition, a performance evaluation of automated systems that attempt to tackle sequence analyses rapidly and reproducibly is generally missing. In order to quantify the accuracy and reproducibility of function assignments on a genome-wide scale, we have re-annotated the entire genome sequence of Chlamydia trachomatis (serovar D), in a collaborative manner. RESULTS: We have encoded all annotations in a structured format to allow further comparison and data exchange and have used a scale that records the different levels of potential annotation errors according to their propensity to propagate in the database due to transitive function assignments. We conclude that genome annotation may entail a considerable amount of errors, ranging from simple typographical errors to complex sequence analysis problems. The most surprising result of this comparative study is that automatic systems might perform as well as the teams of experts annotating genome sequences.

Amino Acid Sequence↗

The chromosomal genome sequence of the spiny sea fan, Muricea muricata (Pallas, 1766) (Malacalcyonacea: Plexauridae) and its associated microbial metagenome sequences.

We present a genome assembly from a Muricea muricata specimen (spiny sea fan; Cnidaria; Anthozoa; Malacalcyonacea; Plexauridae). The genome sequence has a total length of 453.40 megabases. Most of the assembly (98.45%) is scaffolded into 16 chromosomal pseudomolecules. The mitochondrial genome has also been assembled, with a length of 19.29 kilobases. Gene annotation of this assembly by Ensembl identified 52 164 protein-coding genes. From the metagenome data, we recovered five bins, of which three were high-quality MAGs.

Malacalcyonacea↗

The chromosomal genome sequence of the lesser starlet coral, Siderastrea radians (Pallas, 1766) (Scleractinia: Rhizangiidae) and its associated microbial metagenome sequences.

We present a genome assembly from a specimen of Siderastrea radians (lesser starlet coral; Cnidaria; Anthozoa; Scleractinia; Rhizangiidae). The genome sequence has a total length of 807.19 megabases. Most of the assembly (94.17%) is scaffolded into 14 chromosomal pseudomolecules. The mitochondrial genome has also been assembled, with a length of 19.38 kilobases. Gene annotation of this assembly by Ensembl identified 47 051 protein-coding genes. From the metagenome data, we recovered two binned metagenomes assigned to the bacterial phylum Bacteroidota and class Bacteroidia.

Scleractinia↗

The chromosomal genome sequence of the maze coral, Meandrina meandrites (Linnaeus, 1758) (Scleractinia: Meandrinidae) and its associated microbial metagenome sequences.

We present a genome assembly from a specimen of Meandrina meandrites (maze coral; Cnidaria; Anthozoa; Scleractinia; Meandrinidae). The genome sequence has a total length of 551.16 megabases. Most of the assembly (99.25%) is scaffolded into 14 chromosomal pseudomolecules. The mitochondrial genome has also been assembled, with a length of 17.2 kilobases. Gene annotation of this assembly by Ensembl identified 30 464 protein-coding genes. We recovered two bins from the metagenome data.

Meandrina meandrites↗

The complete chloroplast genome sequence of Citrus sinensis (L.) Osbeck var 'Ridge Pineapple': organization and phylogenetic relationships to other angiosperms.

BACKGROUND: The production of Citrus, the largest fruit crop of international economic value, has recently been imperiled due to the introduction of the bacterial disease Citrus canker. No significant improvements have been made to combat this disease by plant breeding and nuclear transgenic approaches. Chloroplast genetic engineering has a number of advantages over nuclear transformation; it not only increases transgene expression but also facilitates transgene containment, which is one of the major impediments for development of transgenic trees. We have sequenced the Citrus chloroplast genome to facilitate genetic improvement of this crop and to assess phylogenetic relationships among major lineages of angiosperms. RESULTS: The complete chloroplast genome sequence of Citrus sinensis is 160,129 bp in length, and contains 133 genes (89 protein-coding, 4 rRNAs and 30 distinct tRNAs). Genome organization is very similar to the inferred ancestral angiosperm chloroplast genome. However, in Citrus the infA gene is absent. The inverted repeat region has expanded to duplicate rps19 and the first 84 amino acids of rpl22. The rpl22 gene in the IRb region has a nonsense mutation resulting in 9 stop codons. This was confirmed by PCR amplification and sequencing using primers that flank the IR/LSC boundaries. Repeat analysis identified 29 direct and inverted repeats 30 bp or longer with a sequence identity > or = 90%. Comparison of protein-coding sequences with expressed sequence tags revealed six putative RNA edits, five of which resulted in non-synonymous modifications in petL, psbH, ycf2 and ndhA. Phylogenetic analyses using maximum parsimony (MP) and maximum likelihood (ML) methods of a dataset composed of 61 protein-coding genes for 30 taxa provide strong support for the monophyly of several major clades of angiosperms, including monocots, eudicots, rosids and asterids. The MP and ML trees are incongruent in three areas: the position of Amborella and Nymphaeales, relationship of the magnoliid genus Calycanthus, and the monophyly of the eurosid I clade. Both MP and ML trees provide strong support for the monophyly of eurosids II and for the placement of Citrus (Sapindales) sister to a clade including the Malvales/Brassicales. CONCLUSION: This is the first complete chloroplast genome sequence for a member of the Rutaceae and Sapindales. Expansion of the inverted repeat region to include rps19 and part of rpl22 and presence of two truncated copies of rpl22 is unusual among sequenced chloroplast genomes. Availability of a complete Citrus chloroplast genome sequence provides valuable information on intergenic spacer regions and endogenous regulatory sequences for chloroplast genetic engineering. Phylogenetic analyses resolve relationships among several major clades of angiosperms and provide strong support for the monophyly of the eurosid II clade and the position of the Sapindales sister to the Brassicales/Malvales.

Ananas↗

Endogenous retroviruses in the human genome sequence.

The human genome contains many endogenous retroviral sequences, and these have been suggested to play important roles in a number of physiological and pathological processes. Can the draft human genome sequences help us to define the role of these elements more closely?

Autoimmune Diseases↗

Detecting and analyzing DNA sequencing errors: toward a higher quality of the Bacillus subtilis genome sequence.

During the determination of a DNA sequence, the introduction of artifactual frameshifts and/or in-frame stop codons in putative genes can lead to misprediction of gene products. Detection of such errors with a method based on protein similarity matching is only possible when related sequences are available in databases. Here, we present a method to detect frameshift errors in DNA sequences that is based on the intrinsic properties of the coding sequences. It combines the results of two analyses, the search for translational initiation/termination sites and the prediction of coding regions. This method was used to screen the complete Bacillus subtilis genome sequence and the regions flanking putative errors were resequenced for verification. This procedure allowed us to correct the sequence and to analyze in detail the nature of the errors. Interestingly, in several cases in-frame termination codons or frameshifts were not sequencing errors but confirmed to be present in the chromosome, indicating that the genes are either nonfunctional (pseudogenes) or subject to regulatory processes such as programmed translational frameshifts. The method can be used for checking the quality of the sequences produced by any prokaryotic genome sequencing project.

Bacillus subtilis↗

Genomic Sequencing in Neonatal Encephalopathy and Suspected Hypoxic-Ischaemic Encephalopathy: A Systematic Review.

BACKGROUND: Neonatal encephalopathy (NE) is a major cause of neonatal mortality and long-term neurological disability. Although hypoxic-ischaemic encephalopathy (HIE) is the most common cause, several genetic disorders may mimic or coexist with hypoxic-ischaemic injury. Next-generation sequencing has emerged as a promising diagnostic tool in this setting. This systematic review evaluated the current evidence on genomic sequencing in NE. MATERIAL AND METHODS: A systematic review was conducted according to PRISMA 2020 guidelines and prospectively registered in PROSPERO. PubMed/MEDLINE, Embase, and Scopus were searched from inception to June 2026. Eligible studies included neonates (≤28 days) with NE, suspected or confirmed HIE, HIE mimics, or unexplained NE who underwent genomic sequencing. Whole-exome sequencing (WES), whole-genome sequencing (WGS), clinical exome sequencing (CES), rapid genomic sequencing, and targeted next-generation sequencing panels were considered. Study quality was assessed using the Newcastle-Ottawa Scale. RESULTS: Seven studies met the inclusion criteria. Considerable heterogeneity was observed regarding patient selection, sequencing strategies, and reported outcomes. Among diagnostic sequencing studies, diagnostic yield ranged from 23.5% to 53.1%. Pathogenic and likely pathogenic variants were identified in genes associated with developmental and epileptic encephalopathies, metabolic disorders, mitochondrial diseases, and neurodevelopmental syndromes, including SCN2A, KCNQ2, CACNA1A, STXBP1, PTPN11, BCOR, MMUT, COQ2, and GBE1. Genomic sequencing frequently refined or changed the initial diagnosis, improved prognostic assessment and genetic counselling, and, in selected cases, guided disease-specific treatment. One study investigated genetic susceptibility to hypoxic-ischaemic injury rather than diagnostic sequencing. CONCLUSIONS: Genomic sequencing provides clinically meaningful diagnoses in a substantial proportion of neonates with unexplained NE or atypical HIE presentations. Current evidence supports integrating genomic sequencing into the diagnostic evaluation of selected infants, although larger prospective studies are needed to define its optimal timing, clinical utility, and cost-effectiveness.

Humans↗

Analysis of human mRNAs with the reference genome sequence reveals potential errors, polymorphisms, and RNA editing.

The NCBI Reference Sequence (RefSeq) project and the NIH Mammalian Gene Collection (MGC) together define a set of approximately 30,000 nonredundant human mRNA sequences with identified coding regions representing 17,000 distinct loci. These high-quality mRNA sequences allow for the identification of transcribed regions in the human genome sequence, and many researchers accept them as the correct representation of each defined gene sequence. Computational comparison of these mRNA sequences and the recently published essentially finished human genome sequence reveals several thousand undocumented nonsynonymous substitution and frame shift discrepancies between the two resources. Additional analysis is undertaken to verify that the euchromatic human genome is sufficiently complete--containing nearly the whole mRNA collection, thus allowing for a comprehensive analysis to be undertaken. Many of the discrepancies will prove to be genuine polymorphisms in the human population, somatic cell genomic variants, or examples of RNA editing. It is observed that the genome sequence variant has significant additional support from other mRNAs and ESTs, almost four times more often than does the mRNA variant, suggesting that the genome sequence is more accurate. In approximately 15% of these cases, there is substantial support for both variants, suggestive of an undocumented polymorphism. An initial screening against a 24-individual genomic DNA diversity panel verified 60% of a small set of potential single nucleotide polymorphisms from which successful results could be obtained. We also find statistical evidence that a few of these discrepancies are due to RNA editing. Overall, these results suggest that the mRNA collections may contain a substantial number of errors. For current and future mRNA collections, it may be prudent to fully reconcile each genome sequence discrepancy, classifying each as a polymorphism, site of RNA editing or somatic cell variation, or genome sequence error.

Computational Biology↗

Targeted reflex RNA sequencing for enhanced variant classification on exome and genome sequencing improves patient outcomes.

RNA sequencing (RNA-seq) has been utilized to provide functional evidence regarding the impact of splicing variants. This study explores the utility of targeted reflex RNA-seq to inform classification of predicted splicing variants identified through clinical exome sequencing (ES) and genome sequencing (GS). A retrospective analysis was conducted on consecutive ES/GS cases completed at a single center in which targeted reflex RNA-seq was performed following identification of eligible variants. There were 131 cases (4.1%) that had at least one RNA-seq eligible variant reported, with eight of these cases having two unique eligible variants. Of the 139 eligible variants, 125 were classified as variants of uncertain significance (VUS). Sixty-four cases had targeted reflex RNA-seq completed with 27 cases having at least one variant reclassified (42.2%). After reclassification, 23 cases had positive results, and two cases had a likely diagnosis of an autosomal recessive condition. Clinical outcomes data regarding positive RNA-seq cases showed that 71% (10/14) had clinical management changes and 43% (6/14) had treatment changes. Incorporation of targeted reflex RNA-seq analysis into the diagnostic pipeline of rare diseases enhances variant classification and resolves uncertainty regarding predicted splice variants, leading to an estimated 1.6% increase in diagnostic yield of clinical ES/GS.

Journal Article↗

Porcine T-cell receptor beta-chain: a genomic sequence covering Dbeta1.1 to Cbeta2 gene segments and the diversity of cDNA expressed in piglets including novel alternative splicing products.

Porcine TCRbeta-chain cDNA clones were isolated from thymic and peripheral blood lymphocytes of piglets. Using these nucleotide sequences, a genomic 18kbp sequence stretch covering Dbeta1 to Cbeta2 gene segments was identified, which revealed that the porcine TCRbeta-chain locus consists of two sets of Dbeta-Jbeta-Cbeta gene groups with each set having a Dbeta gene segment, seven Jbeta gene segments and a down stream Cbeta gene segment composed of four exons. This structure is consistent with other known mammalian TCRbeta-chain loci. With this genomic information, TCRbeta-chain clones from cDNA libraries were analyzed. Sixteen Vbeta gene segments were obtained accompanied by either Dbeta1 or Dbeta2 and by one of the nine Jbeta gene segments. Five different Cbeta cDNA sequences were obtained including four types of Cbeta1 sequences and one type of Cbeta2 sequence. The differences among the Cbeta1 sequences are either allelic polymorphisms or two splice variants, one being a product of exon1 splicing to exon3 (exon2 skipping), and another being an alternative splicing using a splice acceptor site newly discovered inside Cbeta1 exon4. The latter splice acceptor site was also found in human, mouse and horse all giving short cytoplasmic domain with Phe at their C-terminal ends. Other splicing products included trans-splicing of Jbeta2 to Cbeta1, non-functional splicing of two Jbeta gene segments in tandem and a part of Jbeta2.7-Cbeta2 intron to Cbeta2 exon1. Numerous examples of splice variants may suggest the involvement of splicing in generating TCRbeta-chain functional diversity.

Alternative Splicing↗

The cost and cost trajectory of genome sequencing and bioinformatics analysis for Indigenous children with suspected rare diseases.

PURPOSE: Indigenous peoples are underrepresented in reference genome libraries. Consequently, rare disease diagnosis may require bespoke bioinformatics analyses of genome sequences. Establishing diagnostic cost is crucial to support policy development for equitable diagnosis of rare diseases. We estimated the cost and cost trajectory of diagnostic genome sequencing and bioinformatics for Indigenous participants with suspected rare diseases. METHODS: We conducted a microcosting study of Indigenous children and their families receiving genome sequencing through Canada's Silent Genomes Project. Invoice data informed the costs of genome sequencing. We conducted a time-and-motion study for bioinformatics analyses, including labor, computing, and data storage costs. RESULTS: With standard bioinformatics, costs ranged from C$3645 (SD: 455) for singletons to C$7402 (SD: 566) for trios. With advanced, bespoke bioinformatics, costs ranged from C$5344 (SD: 634) for singletons to C$9760 (SD: 822) for trios. Genome sequencing was a primary cost driver; however, sequencing costs decreased by 61% over 4 years. Bioinformatics costs ranged from 21.3% to 58.3% of the total costs. The time required for bioinformatics ranged from 71 hours to 215 hours for standard and advanced analyses, respectively. CONCLUSION: Genome sequencing costs decreased over time. Bioinformatics is a significant cost driver, particularly for bespoke analyses arising from nonrepresentative reference libraries.

Humans↗

Genomic sequence of the RA27/3 vaccine strain of rubella virus.

The sequence of the genome of the RA27/3 vaccine strain of rubella virus (RUB) was determined. In the process, several discrepancies between the previously reported genomic sequences of two wild RUB strains (Therien and M33) were resolved. The genomes of all three strains contain 9762 nucleotides (nts), exclusive of the 3' poly A tract. In all three strains, the genome contains (5' to 3'), a 40 nt 5' untranslated region (UTR), an open reading frame (ORF) of 6348 nts that encodes nonstructural proteins, a 123 nt UTR between the two genomic ORFs, a 3189 nt ORF that encodes the structural proteins, and a 62 nt 3' UTR. The 5' end of the subgenomic RNA was found to correspond to a uridine residue at nt 6436 of the genomic RNA. At the nucleotide level, the sequence of the three strains varied by 1.0 to 2.8%, while at the amino acid level, the sequence varied by 1.1 to 2.4% over both ORFs. The RA27/3 sequence will be of use in identification of the determinants of its attenuation, in vaccine production control and in development of second generation RUB vaccines based on recombinant DNA technology.

Amino Acid Sequence↗

Completion of the genome sequence of Brucella abortus and comparison to the highly similar genomes of Brucella melitensis and Brucella suis.

Brucellosis is a worldwide disease of humans and livestock that is caused by a number of very closely related classical Brucella species in the alpha-2 subdivision of the Proteobacteria. We report the complete genome sequence of Brucella abortus field isolate 9-941 and compare it to those of Brucella suis 1330 and Brucella melitensis 16 M. The genomes of these Brucella species are strikingly similar, with nearly identical genetic content and gene organization. However, a number of insertion-deletion events and several polymorphic regions encoding putative outer membrane proteins were identified among the genomes. Several fragments previously identified as unique to either B. suis or B. melitensis were present in the B. abortus genome. Even though several fragments were shared between only B. abortus and B. suis, B. abortus shared more fragments and had fewer nucleotide polymorphisms with B. melitensis than B. suis. The complete genomic sequence of B. abortus provides an important resource for further investigations into determinants of the pathogenicity and virulence phenotypes of these bacteria.

Bacterial Proteins↗

Evaluating the return of additional findings from the 100,000 Genomes Project: A mixed-methods study exploring participant experiences of receiving secondary findings from genomic sequencing.

PURPOSE: The 100,000 Genomes Project participants could consent to receive additional findings (AFs) for variants associated with susceptibility to cancer and familial hypercholesterolemia. Here, we evaluate stakeholder experiences to inform clinical practice. METHODS: Mixed-methods study conducted at 18 sites across England that comprised a cross-sectional survey and interviews with participants who received a positive AF (PAF) and interviews with participants who had no AFs (NAF). RESULTS: There were 146 surveys followed by 35 interviews with PAF participants and 29 interviews with NAF participants. Surveys found that PAF results were seen as useful and would influence health management (82%). Most (90%) had shared their result with family members. Experiences differed by PAF type; cancer PAF participants were often initially shocked and anxious and found telling family members challenging compared with participants with a familial hypercholesterolemia PAF. Although most experiences of NAF results were positive, some misunderstandings were identified. Participants supported returning AFs when offering genome sequencing. CONCLUSION: Patient experiences of receiving AFs were primarily positive, and there is support for offering AFs routinely. Considerations for offering AFs in clinical practice include adapting approaches tailored to individual conditions and greater support for people with a NAF result.

Humans↗

Genome sequencing and annotation: an overview.

Many microbial genome sequences have been determined, and more new genome projects are ongoing. Shotgun sequencing of randomly cloned short pieces of genomic DNA can provide a simple way of determining whole genome sequences. This process requires sequencing of many fragments, compilation of the separate sequences into one contiguous sequence, and careful editing of the assembled sequence. The genes present on the microbial genome are then predicted using clues derived from typical gene features, such as codon usage, ribosomal binding sequences, and bacterial initiation codons. Function of genes is predicted by homology searches performed against either public or well-established protein databases. This chapter discusses each of these stages in a genome-sequencing project.

Amino Acid Sequence↗