Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

Whole-genome sequencing of adenovirus 41 directly from wastewater using nested overlapping PCR and MinION.

Human adenovirus F41 (HAdV-F41) is one of the leading causes of children's acute gastroenteritis and was recently linked to an outbreak of severe acute hepatitis of unknown etiology among children during 2021 to 2022. While most evidence is based on clinical data, wastewater-based epidemiology offers a community-level approach to monitoring circulating strains and enhancing outbreak preparedness. In this study, we developed an overlapping amplicon-based whole-genome sequencing approach to directly detect HAdV-F41 from archived wastewater samples, using nested PCR with 13 primer sets. Archived wastewater samples were collected between 2021 and 2022 from three treatment plants in Seattle, USA. The viral load ranged from 1.2 × 103 to 8.4 × 103 genome copies per liter. The Oxford Nanopore platform was used for whole-genome sequencing. Complete or partial (>84%) HAdV-F41 genomes were recovered from wastewater samples, with mean coverage depths ranging from 10³ to 10⁵. The consensus sequences showed more than 99% similarity to reference genomes in the NCBI database. The phylogenetic analysis revealed that 2 sequences clustered within lineage 2a and 11 within lineage 2b, reflecting that at least two sub-lineages were circulating in the community at that time. Our results demonstrate that the overlapping amplicon-based whole-genome sequencing approach using the Oxford Nanopore platform reliably recovers HAdV-F41 genomes from wastewater. This method offers high-resolution genomic surveillance of circulating, clinically relevant HAdV-F41, supporting wastewater-based epidemiology as a valuable tool for detecting emerging variants and strengthening the early warning system for future disease outbreaks.IMPORTANCEHuman adenovirus F41 is a primary cause of childhood gastroenteritis and has been linked to recent outbreaks of severe acute hepatitis in children, yet community-level genomic surveillance of this virus remains limited. This study shows that wastewater can be used to recover nearly complete HAdV-F41 genomes through a targeted overlapping-amplicon sequencing strategy on the Oxford Nanopore platform. By applying this method to archived wastewater samples, we detected the simultaneous circulation of multiple viral lineages in a large city. These findings extend wastewater-based epidemiology beyond SARS-CoV-2 and emphasize its importance for monitoring clinically significant enteric viruses. The method described here offers a scalable tool for tracking viral evolution in communities and enhancing early warning systems for future outbreaks.

Wastewater↗

Markov encoding for detecting signals in genomic sequences.

We present a technique to encode the inputs to neural networks for the detection of signals in genomic sequences. The encoding is based on lower-order Markov models which incorporate known biological characteristics in genomic sequences. The neural networks then learn intrinsic higher-order dependencies of nucleotides at the signal sites. We demonstrate the efficacy of the Markov encoding method in the detection of three genomic signals, namely, splice sites, transcription start sites, and translation initiation sites.

Base Sequence↗

The biology of enhancer-dependent transcriptional regulation in bacteria: insights from genome sequences.

The bacterial transcription factor sigma(N) (sigma-N, sigma-54, RpoN) confers upon RNA polymerase (RNAP) properties distinct from those of the major house-keeping form of RNAP, which contains sigma(70) (sigma-70, RpoD). Transcription by RNAP containing sigma(N) is subject to enhancer-dependent regulation. Far from being an 'oddity' or 'exception to the rule', the occurrence of sigma(N) in the genome sequences of such diverse bacteria as Aquifex aeolicus, Bacillus subtilis, Chlamydia spp. and Borrelia burgdorferi argues for its biological importance. The availability of complete genome sequences of several (eu)bacteria offers an opportunity to extend our understanding of this special form of transcriptional regulation. By scanning their genome sequences, new functions have been predicted for enhancer-dependent transcription in A. aeolicus, Chlamydia trachomatis, Escherichia coli, Treponema pallidum and B. burgdorferi.

Bacteria↗

A preliminary gene map for the Van der Woude syndrome critical region derived from 900 kb of genomic sequence at 1q32-q41.

Van der Woude syndrome (VWS) is a common form of syndromic cleft lip and palate and accounts for approximately 2% of all cleft lip and palate cases. Distinguishing characteristics include cleft lip with or without cleft palate, isolated cleft palate, bilateral lip pits, hypodontia, normal intelligence, and an autosomal-dominant mode of transmission with a high degree of penetrance. Previously, the VWS locus was mapped to a 1.6-cM region in 1q32-q41 between D1S491 and D1S205, and a 4.4-Mb contig of YAC clones of this region was constructed. In the current investigation, gene-based and anonymous STSs were developed from the existing physical map and were then used to construct a contig of sequence-ready bacterial clones across the entire VWS critical region. All STSs and BAC clones were shared with the Sanger Centre, which developed a contig of PAC clones over the same region. A subset of 11 clones from both contigs was selected for high-throughput sequence analysis across the approximately 1.1-Mb region; all but two of these clones have been sequenced completely. Over 900 kb of genomic sequence, including the 350-kb VWS critical region, were analyzed and revealed novel polymorphisms, including an 8-kb deletion/insertion, and revealed 4 known genes, 11 novel genes, 9 putative genes, and 3 psuedogenes. The positional candidates LAMB3, G0S2, HIRF6, and HSD11 were excluded as the VWS gene by mutation analysis. A preliminary gene map for the VWS critical region is as follows: [see text] 41-TEL. The data provided here will help lead to the identification of the VWS gene, and this study provides a model for how laboratories that have a regional interest in the human genome can contribute to the sequencing efforts of the entire human genome.

Animals↗

Complete genomic sequence of bacteriophage ul36: demonstration of phage heterogeneity within the P335 quasi-species of lactococcal phages.

The complete genomic sequence of the Lactococcus lactis virulent phage ul36 belonging to P335 lactococcal phage species was determined and analyzed. The genomic sequence of this lactococcal phage contained 36,798 bp with an overall G+C content of 35.8 mol %. Fifty-nine open reading frames (ORFs) of more than 40 codons were found. N-terminal sequencing of phage structural proteins as well as bioinformatic analysis led to the attribution of a function to 24 ORFs (41%). A lysogeny module was found within the genome of this virulent phage. The putative integrase gene seems to be the product of a horizontal transfer because it is more closely related to Streptococcus pyogenes phages than it is to L. lactis phages. Comparative genome analysis with six complete genomes of temperate P335-like phages confirmed the heterogeneity among phages of P335 species. A dUTPase gene is the only conserved gene among all P335 phages analyzed as well as the phage BK5-T. A genetic relationship between P335 phages and the phage-type of the BK5-T species was established. Thus, we proposed that phage BK5-T be included within the P335 species and thereby reducing the number of lactococcal phage species to 11.

Bacteriophages↗

Scalable approaches for functional analyses of whole-genome sequencing non-coding variants.

Non-coding genetic variants outside of protein-coding genome regions play an important role in genetic and epigenetic regulation. It has become increasingly important to understand their roles, as non-coding variants often make up the majority of top findings of genome-wide association studies (GWAS). In addition, the growing popularity of disease-specific whole-genome sequencing (WGS) efforts expands the library of and offers unique opportunities for investigating both common and rare non-coding variants, which are typically not detected in more limited GWAS approaches. However, the sheer size and breadth of WGS data introduce additional challenges to predicting functional impacts in terms of data analysis and interpretation. This review focuses on the recent approaches developed for efficient, at-scale annotation and prioritization of non-coding variants uncovered in WGS analyses. In particular, we review the latest scalable annotation tools, databases and functional genomic resources for interpreting the variant findings from WGS based on both experimental data and in silico predictive annotations. We also review machine learning-based predictive models for variant scoring and prioritization. We conclude with a discussion of future research directions which will enhance the data and tools necessary for the effective functional analyses of variants identified by WGS to improve our understanding of disease etiology.

Genome-Wide Association Study↗

Protein families and TRIBES in genome sequence space.

Accurate detection of protein families allows assignment of protein function and the analysis of functional diversity in complete genomes. Recently, we presented a novel algorithm called TribeMCL for the detection of protein families that is both accurate and efficient. This method allows family analysis to be carried out on a very large scale. Using TribeMCL, we have generated a resource called TRIBES that contains protein family information, comprising annotations, protein sequence alignments and phylogenetic distributions describing 311 257 proteins from 83 completely sequenced genomes. The analysis of at least 60 934 detected protein families reveals that, with the essential families excluded, paralogy levels are similar between prokaryotes, irrespective of genome size. The number of essential families is estimated to be between 366 and 426. We also show that the currently known space of protein families is scale free and discuss the implications of this distribution. In addition, we show that smaller families are often formed by shorter proteins and discuss the reasons for this intriguing pattern. Finally, we analyse the functional diversity of protein families in entire genome sequences. The TRIBES protein family resource is accessible at http://www.ebi.ac.uk/research/cgg/tribes/.

Algorithms↗

Identification by bacterial expression and functional reconstitution of the yeast genomic sequence encoding the mitochondrial dicarboxylate carrier protein.

The inner membranes of mitochondria contain a family of transport proteins of related sequence and structure. The DNA sequence of the genome of Saccharomyces cerevisiae encodes at least 35 members of this family. Three of them can be recognised as known isoforms of the ADP-ATP translocase and two others as the phosphate and citrate carriers. The transport functions of the remainder cannot be identified with certainty. One of them, encoded on yeast chromosome xii, shows a fairly close sequence relationship to the known sequence of the bovine mitochondrial oxoglutarate-malate carrier. The yeast protein has been obtained by over-expression in Escherichia coli, reconstituted into phospholipid vesicles and shown to have transport properties characteristic of the mitochondrial carrier for dicarboxylate ions, such as malate, and also phosphate, previously biochemically characterised, but not sequenced, from both mammalian and yeast mitochondria. This is the first example of the biochemical identification of an unknown membrane protein encoded in the yeast genome since the completion of the genomic sequence.

Carrier Proteins↗

Full genome sequence of peste des petits ruminants virus, a member of the Morbillivirus genus.

Peste des petits ruminants virus (PPRV) causes an acute febrile illness in small ruminant species, mostly sheep and goats. PPRV is a member of the Morbillivirus genus which includes measles, rinderpest (cattle plague), canine distemper, phocine distemper and the morbilliviruses found in whales, porpoises and dolphins. Full length genome sequences for these morbilliviruses are available and reverse genetic rescue systems have been developed for the viruses of terrestrial mammals, with the exception of PPRV. This paper presents the first published full length genome sequence for PPRV. The genome was found to be consistent with the rule-of-six and open reading frames (ORFs) were identified that encoded the eight proteins characteristic of morbilliviruses. At the nucleotide (nt) level, the full length genome of PPRV was most similar to that of rinderpest, the other ruminant morbillivirus. However, at the protein level five of the six structural proteins and the V protein showed a greater similarity to the dolphin morbillivirus (DMV) while only the C and L proteins showed a high relationship to rinderpest.

Base Sequence↗

Human and mouse alpha-synuclein genes: comparative genomic sequence analysis and identification of a novel gene regulatory element.

The human alpha-synuclein gene (SNCA) encodes a presynaptic nerve terminal protein that was originally identified as a precursor of the non-beta-amyloid component of Alzheimer's disease plaques. More recently, mutations in SNCA have been identified in some cases of familial Parkinson's disease, presenting numerous new areas of investigation for this important disease. Molecular studies would benefit from detailed information about the long-range sequence context of SNCA. To that end, we have established the complete genomic sequence of the chromosomal regions containing the human and mouse alpha-synuclein genes, with the objective of using the resulting sequence information to identify conserved regions of biological importance through comparative sequence analysis. These efforts have yielded approximately 146 and approximately 119 kb of high-accuracy human and mouse genomic sequence, respectively, revealing the precise genetic architecture of the alpha-synuclein gene in both species. A simple repeat element upstream of SNCA/Snca has been identified and shown to be necessary for normal expression in transient transfection assays using a luciferase reporter construct. Together, these studies provide valuable data that should facilitate more detailed analysis of this medically important gene.

Animals↗

Complete genome sequence of the Lactococcus lactis temperate phage phiLC3: comparative analysis of phiLC3 and its relatives in lactococci and streptococci.

Complete genome sequencing of the P335 temperate Lactococcus lactis bacteriophage phiLC3 (32, 172 bp) revealed fifty-one open reading frames (ORFs). Four ORFs did not show any homology to other proteins in the database and twenty-one ORFs were assigned a putative biological function. phiLC3 contained a unique replication module and orf201 was identified as the putative replication initiator protein-encoding gene. phiLC3 was closely related to the L. lactis r1t phage (73% DNA identity). Similarity was also shared with other lactococcal P335 phages and the Streptococcus pyogenes prophages 370.3, 8232.4 and 315.5 over the non-structural genes and the genes involved in DNA packaging/phage morphogenesis, respectively. phiLC3 contained small homologous regions distributed among lactococcal phages suggesting that these regions might be involved in mediating genetic exchange. Two regions of 30 and 32 bp were conserved among the streptococcal and lactococcal r1t-like phages. These two regions, as well as other homologous regions, were located at mosaic borders and close to putative transcriptional terminators indicating that such regions together might attract recombination. The conserved regions found among lactococcal and streptococcal phages might be used for identification of phages/prophages/prophage remnants in their hosts.

Base Sequence↗

A computer simulation analysis of the accuracy of partial genome sequencing and restriction fragment analysis in the reconstruction of phylogenetic relationships.

Partial genome sequencing (PGS) and restriction fragment analysis (RFA) are used frequently in molecular epidemiologic investigations. The relative accuracy of PGS and RFA in phylogenetic reconstruction has not been assessed. In this study, 32 model phylogenetic trees with 16 extant lineages were generated, for which DNA sequences were simulated under varying conditions of genome length, nucleotide substitution rate, and between-site substitution rate variation. Genotyping using PGS and RFA was simulated. The effect of tree structure (stemminess, imbalance, lineage variation) on the accuracy of phylogenetic reconstruction (topological and branch length similarity) was evaluated. Overall, PGS was more accurate than RFA. The accuracy of PGS increased with increasing sequence length. The accuracy of RFA increased with the number of restriction enzymes used. In fragment size comparison, the Dice and Nei-Li algorithms differed little, with both more accurate than the Fragment Size Distribution algorithm. For RFA, higher tree stemminess and longer genome length were associated with higher topological accuracy, whereas lower tree stemminess and lower substitution rates were associated with higher branch length accuracy. For PGS, lower tree imbalance was associated with higher topological accuracy, whereas lower tree stemminess, higher substitution rate, and lower between-site substitution rate variation were associated with higher branch length accuracy. RFA had higher topological accuracy than PGS only for the shortest sequence length (200 bps) at a low substitution rate, high tree stemminess, and long genome length. PGS had equal or higher accuracy in branch length reconstruction than RFA under all conditions investigated. Thus, partial genome sequencing is recommended over restriction fragment analysis for conditions within the parameter space examined.

Computational Biology↗

[Annotation of complete genomic sequence of 3p24-p25 478 kb of human DNA].

OBJECTIVE: To annotate the human genome 3p24-p25 478 kb complete sequence. METHODS: The protein-coding genes in the genomic sequence were identified by using ab initio gene finding, homology-based similarity database searching and all or partial mRNA aligning with genomic sequence, and the content feature of the genomic sequence were analyzed by using EMBOSS package. RESULTS: Two known genes SLC6A1 and SLC6A11 were identified; as well as the GC content of this genomic sequence was 47% and 3 putative CpG islands were predicted in the genomic sequence, located in 130,685-131,516 bp, 307,090-307,870 bp and 415,585-416,308 bp, respectively. CONCLUSIONS: The methods, as mentioned above, might be used for annotating the biological information in the genomic sequence, such as gene structure, GC content, CpG island.

Base Sequence↗

Genome sequence of a serotype M3 strain of group A Streptococcus: phage-encoded toxins, the high-virulence phenotype, and clone emergence.

Genome sequences are available for many bacterial strains, but there has been little progress in using these data to understand the molecular basis of pathogen emergence and differences in strain virulence. Serotype M3 strains of group A Streptococcus (GAS) are a common cause of severe invasive infections with unusually high rates of morbidity and mortality. To gain insight into the molecular basis of this high-virulence phenotype, we sequenced the genome of strain MGAS315, an organism isolated from a patient with streptococcal toxic shock syndrome. The genome is composed of 1,900,521 bp, and it shares approximately 1.7 Mb of related genetic material with genomes of serotype M1 and M18 strains. Phage-like elements account for the great majority of variation in gene content relative to the sequenced M1 and M18 strains. Recombination produces chimeric phages and strains with previously uncharacterized arrays of virulence factor genes. Strain MGAS315 has phage genes that encode proteins likely to contribute to pathogenesis, such as streptococcal pyrogenic exotoxin A (SpeA) and SpeK, streptococcal superantigen (SSA), and a previously uncharacterized phospholipase A(2) (designated Sla). Infected humans had anti-SpeK, -SSA, and -Sla antibodies, indicating that these GAS proteins are made in vivo. SpeK and SSA were pyrogenic and toxic for rabbits. Serotype M3 strains with the phage-encoded speK and sla genes increased dramatically in frequency late in the 20th century, commensurate with the rise in invasive disease caused by M3 organisms. Taken together, the results show that phage-mediated recombination has played a critical role in the emergence of a new, unusually virulent clone of serotype M3 GAS.

Amino Acid Sequence↗

[Complete genome sequence analysis of the Hantavirus Z10 strain].

OBJECTIVE: Study on the complete genome sequence of Hantavirus Z10 strain which has been applied for inactivated vaccine production in China, to assess its molecular characteristics and the diversity with other hantaviruses. METHODS: The total RNA were prepared from Z10 virus infected cells and the RT-PCR products was cloned into T vector, sequenced and analyzed by using DNASTAR software. RESULTS: The Z10 complete genome, L segment is 6,553, M segment is 3,615, S segment is 1,701 nucleotides in length, with a single open reading frame encoding 2,151, 1,135, 429 amino acids respectively. Sequence homology comparison showed that the 3 segment nucleotide of Z10 strain were close to HTN type virus, but only 83.6-87.4% homology with other HTN viruses at the nucleotide level. The phylogenetic analysis was made on their nucleotide and amino acid sequences. CONCLUSION: The results firstly demonstrates that Z10 strain is a new subtype of the Hantaan(HTN) type.

Amino Acid Sequence↗

Shark (Scyliorhinus torazame) metallothionein: cDNA cloning, genomic sequence, and expression analysis.

Novel metallothionein (MT) complementary DNA and genomic sequences were isolated from a cartilaginous shark species, Scyliorhinus torazame. The full-length open reading frame (ORF) of shark MT cDNA encoded 68 amino acids with a high cysteine content (29%). The genomic ORF sequence (932 bp) of shark MT isolated by polymerase chain reaction (PCR) comprised 3 exons with 2 interventing introns. Shark MT sequence shared many conserved features with other vertebrate MTs: overall amino acid identities of shark MT ranged from 47% to 57% with fish MTs, and 41% to 62% with mammalian MTs. However, in addition to these conserved characteristics, shark MT sequence exhibited some unique characteristics. It contained 4 extra amino acids (Lys-Ala-Gly-Arg) at the end of the beta-domain, which have not been reported in any other vertebrate MTs. The last amino acid residue at the C-terminus was Ser, which also has not been reported in fish and mammalian MTs. The MT messenger RNA levels in shark liver and kidney, assessed by semiquantitative reverse transcriptase PCR and RNA blot hybridization, were significantly affected by experimental exposures to heavy metals (cadmium, copper, and zinc). Generally, the transcriptional activation of shark MT gene was dependent on the dose (0-10 mg/kg body weight for injection and 0-20 microM for immersion) and duration (1-10 days); zinc was a more potent inducer than copper and cadmium.

Amino Acid Sequence↗

Deciphering Campylobacter jejuni cell surface interactions from the genome sequence.

The completion of the Campylobacter jejuni genome sequence is a landmark in Campylobacter research. Discoveries directly arising from these data include the identification of a capsular polysaccharide, extensive capacity for phase variable gene expression and lipo-oligosaccharide structural phase variation. The recent identification of a unique system of general protein glycosylation in C. jejuni, a C. jejuni protein that is translocated into eukaryotic cells, and plasmid-encoded components of a putative type IV secretion system are likely to be significant in terms of the host-pathogen interaction.

Antigenic Variation↗

Genome sequences of Chlamydia trachomatis MoPn and Chlamydia pneumoniae AR39.

The genome sequences of Chlamydia trachomatis mouse pneumonitis (MoPn) strain Nigg (1 069 412 nt) and Chlamydia pneumoniae strain AR39 (1 229 853 nt) were determined using a random shotgun strategy. The MoPn genome exhibited a general conservation of gene order and content with the previously sequenced C.trachomatis serovar D. Differences between C.trachomatis strains were focused on an approximately 50 kb 'plasticity zone' near the termination origins. In this region MoPn contained three copies of a novel gene encoding a >3000 amino acid toxin homologous to a predicted toxin from Escherichia coli O157:H7 but had apparently lost the tryptophan biosyntheis genes found in serovar D in this region. The C. pneumoniae AR39 chromosome was >99.9% identical to the previously sequenced C.pneumoniae CWL029 genome, however, comparative analysis identified an invertible DNA segment upstream of the uridine kinase gene which was in different orientations in the two genomes. AR39 also contained a novel 4524 nt circular single-stranded (ss)DNA bacteriophage, the first time a virus has been reported infecting C. pneumoniae. Although the chlamydial genomes were highly conserved, there were intriguing differences in key nucleotide salvage pathways: C.pneumoniae has a uridine kinase gene for dUTP production, MoPn has a uracil phosphororibosyl transferase, while C.trachomatis serovar D contains neither gene. Chromosomal comparison revealed that there had been multiple large inversion events since the species divergence of C.trachomatis and C.pneumoniae, apparently oriented around the axis of the origin of replication and the termination region. The striking synteny of the Chlamydia genomes and prevalence of tandemly duplicated genes are evidence of minimal chromosome rearrangement and foreign gene uptake, presumably owing to the ecological isolation of the obligate intracellular parasites. In the absence of genetic analysis, comparative genomics will continue to provide insight into the virulence mechanisms of these important human pathogens.

Animals↗