Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Comparative genome sequence analysis of the Bpa/Str region in mouse and Man.

The progress of human and mouse genome sequencing programs presages the possibility of systematic cross-species comparison of the two genomes as a powerful tool for gene and regulatory element identification. As the opportunities to perform comparative sequence analysis emerge, it is important to develop parameters for such analyses and to examine the outcomes of cross-species comparison. Our analysis used gene prediction and a database search of 430 kb of genomic sequence covering the Bpa/Str region of the mouse X chromosome, and 745 kb of genomic sequence from the homologous human X chromosome region. We identified 11 genes in mouse and 13 genes and two pseudogenes in human. In addition, we compared the mouse and human sequences using pairwise alignment and searches for evolutionary conserved regions (ECRs) exceeding a defined threshold of sequence identity. This approach aided the identification of at least four further putative conserved genes in the region. Comparative sequencing revealed that this region is a mosaic in evolutionary terms, with considerably more rearrangement between the two species than realized previously from comparative mapping studies. Surprisingly, this region showed an extremely high LINE and low SINE content, low G+C content, and yet a relatively high gene density, in contrast to the low gene density usually associated with such regions.

3-Hydroxysteroid Dehydrogenases↗

Mauve: multiple alignment of conserved genomic sequence with rearrangements.

As genomes evolve, they undergo large-scale evolutionary processes that present a challenge to sequence comparison not posed by short sequences. Recombination causes frequent genome rearrangements, horizontal transfer introduces new sequences into bacterial chromosomes, and deletions remove segments of the genome. Consequently, each genome is a mosaic of unique lineage-specific segments, regions shared with a subset of other genomes and segments conserved among all the genomes under consideration. Furthermore, the linear order of these segments may be shuffled among genomes. We present methods for identification and alignment of conserved genomic DNA in the presence of rearrangements and horizontal transfer. Our methods have been implemented in a software package called Mauve. Mauve has been applied to align nine enterobacterial genomes and to determine global rearrangement structure in three mammalian genomes. We have evaluated the quality of Mauve alignments and drawn comparison to other methods through extensive simulations of genome evolution.

Chromosomes, Bacterial↗

A deep-coverage tomato BAC library and prospects toward development of an STC framework for genome sequencing.

Recently a new strategy using BAC end sequences as sequence-tagged connectors (STCs) was proposed for whole-genome sequencing projects. In this study, we present the construction and detailed characterization of a 15.0 haploid genome equivalent BAC library for the cultivated tomato, Lycopersicon esculentum cv. Heinz 1706. The library contains 129,024 clones with an average insert size of 117.5 kb and a chloroplast content of 1.11%. BAC end sequences from 1490 ends were generated and analyzed as a preliminary evaluation for using this library to develop an STC framework to sequence the tomato genome. A total of 1205 BAC end sequences (80.9%) were obtained, with an average length of 360 high-quality bases, and were searched against the GenBank database. Using a cutoff expectation value of <10(-6), and combining the results from BLASTN, BLASTX, and TBLASTX searches, 24.3% of the BAC end sequences were similar to known sequences, of which almost half (48.7%) share sequence similarities to retrotransposons and 7% to known genes. Some of the transposable element sequences were the first reported in tomato, such as sequences similar to maize transposon Activator (Ac) ORF and tobacco pararetrovirus-like sequences. Interestingly, there were no BAC end sequences similar to the highly repeated TGRI and TGRII elements. However, the majority (70.3%) of STCs did not share significant sequence similarities to any sequences in GenBank at either the DNA or predicted protein levels, indicating that a large portion of the tomato genome is still unknown. Our data demonstrate that this BAC library is suitable for developing an STC database to sequence the tomato genome. The advantages of developing an STC framework for whole-genome sequencing of tomato are discussed.

Chromosomes, Bacterial↗

First complete genome sequence of two Staphylococcus epidermidis bacteriophages.

Staphylococcus epidermidis is an important opportunistic pathogen causing nosocomial infections and is often associated with infections in patients with implanted prosthetic devices. A number of virulence determinants have been identified in S. epidermidis, which are typically acquired through horizontal gene transfer. Due to the high recombination potential, bacteriophages play an important role in these transfer events. Knowledge of phage genome sequences provides insights into phage-host biology and evolution. We present the complete genome sequence and a molecular characterization of two S. epidermidis phages, phiPH15 (PH15) and phiCNPH82 (CNPH82). Both phages belonged to the Siphoviridae family and produced stable lysogens. The PH15 and CNPH82 genomes displayed high sequence homology; however, our analyses also revealed important functional differences. The PH15 genome contained two introns, and in vivo splicing of phage mRNAs was demonstrated for both introns. Secondary structures for both introns were also predicted and showed high similarity to those of Streptococcus thermophilus phage 2972 introns. An additional finding was differential superinfection inhibition between the two phages that corresponded with differences in nucleotide sequence and overall gene content within the lysogeny module. We conducted phylogenetic analyses on all known Siphoviridae, which showed PH15 and CNPH82 clustering with Staphylococcus aureus, creating a novel clade within the S. aureus group and providing a higher overall resolution of the siphophage branch of the phage proteomic tree than previous studies. Until now, no S. epidermidis phage genome sequences have been reported in the literature, and thus this study represents the first complete genomic and molecular description of two S. epidermidis phages.

Base Sequence↗

Genome sequence of the dissimilatory metal ion-reducing bacterium Shewanella oneidensis.

Shewanella oneidensis is an important model organism for bioremediation studies because of its diverse respiratory capabilities, conferred in part by multicomponent, branched electron transport systems. Here we report the sequencing of the S. oneidensis genome, which consists of a 4,969,803-base pair circular chromosome with 4,758 predicted protein-encoding open reading frames (CDS) and a 161,613-base pair plasmid with 173 CDSs. We identified the first Shewanella lambda-like phage, providing a potential tool for further genome engineering. Genome analysis revealed 39 c-type cytochromes, including 32 previously unidentified in S. oneidensis, and a novel periplasmic [Fe] hydrogenase, which are integral members of the electron transport system. This genome sequence represents a critical step in the elucidation of the pathways for reduction (and bioremediation) of pollutants such as uranium (U) and chromium (Cr), and offers a starting point for defining this organism's complex electron transport systems and metal ion-reducing capabilities.

Amino Acid Sequence↗

The complete genome sequence of Perina nuda picorna-like virus, an insect-infecting RNA virus with a genome organization similar to that of the mammalian picornaviruses.

Perina nuda picorna-like virus (PnPV) is an insect-infecting RNA virus with morphological and physicochemical characters similar to the Picornaviridae. In this article, we determine the complete genome sequence and analyze the gene organization of PnPV. The genome of PnPV consists of 9476 nucleotides (nts) excluding the poly(A) tail and contains a single large open reading frame (ORF) of 8958 nts (2986 codons) flanked by 473 and 45 nt noncoding regions on the 5' and 3' ends, respectively. Northern blotting did not detect the presence of any subgenomic RNA. The PnPV genome codes for four structural proteins (CP1-4), and determination of their N-terminal sequences by Edman degradation, showed that all four are located in the 5' region of the genome. The 3' part of the PnPV genome contains the consensus sequence motifs for picornavirus RNA helicase, cysteine protease, and RNA-dependent RNA polymerase (RdRp) in that order from the 5' to the 3' end. In all of these characters, the genome organization of PnPV resembles the mammalian picornaviruses and two other insect picorna-like viruses, infectious flacherie virus (IFV) of the silkworm and Sacbrood virus (SBV) of the honeybee. In a phylogenetic tree based on the eight conserved domains in the RdRp sequence, PnPV formed a separate cluster with IFV and SBV, which suggests that these three insect picorna-like viruses might constitute a novel group of insect-infecting RNA viruses.

Amino Acid Sequence↗

Comparative analysis of the complete plastid genome sequence of the red alga Gracilaria tenuistipitata var. liui provides insights into the evolution of rhodoplasts and their relationship to other plastids.

We sequenced to completion the circular plastid genome of the red alga Gracilaria tenuistipitata var. liui. This is the first plastid genome sequence from the subclass Florideophycidae (Rhodophyta). The genome is composed of 183,883 bp and contains 238 predicted genes, including a single copy of the ribosomal RNA operon. Comparisons with the plastid genome of Porphyra pupurea reveal strong conservation of gene content and order, but we found major genomic rearrangements and the presence of coding regions that are specific to Gracilaria. Phylogenetic analysis of a data set of 41 concatenated proteins from 23 plastid and two cyanobacterial genomes support red algal plastid monophyly and a specific evolutionary relationship between the Florideophycidae and the Bangiales. Gracilaria maintains a surprisingly ancient gene content in its plastid genome and, together with other Rhodophyta, contains the most complete repertoire of plastid genes known in photosynthetic eukaryotes.

Consensus Sequence↗

Detecting protein function and protein-protein interactions from genome sequences.

A computational method is proposed for inferring protein interactions from genome sequences on the basis of the observation that some pairs of interacting proteins have homologs in another organism fused into a single protein chain. Searching sequences from many genomes revealed 6809 such putative protein-protein interactions in Escherichia coli and 45,502 in yeast. Many members of these pairs were confirmed as functionally related; computational filtering further enriches for interactions. Some proteins have links to several other proteins; these coupled links appear to represent functional interactions such as complexes or pathways. Experimentally confirmed interacting pairs are documented in a Database of Interacting Proteins.

Amino Acid Sequence↗

The Arabidopsis root transcriptome by serial analysis of gene expression. Gene identification using the genome sequence.

Large-scale identification of genes expressed in roots of the model plant Arabidopsis was performed by serial analysis of gene expression (SAGE), on a total of 144,083 sequenced tags, representing at least 15,964 different mRNAs. For tag to gene assignment, we developed a computational approach based on 26,620 genes annotated from the complete sequence of the genome. The procedure selected warrants the identification of the genes corresponding to the majority of the tags found experimentally, with a high level of reliability, and provides a reference database for SAGE studies in Arabidopsis. This new resource allowed us to characterize the expression of more than 3,000 genes, for which there is no expressed sequence tag (EST) or cDNA in the databases. Moreover, 85% of the tags were specific for one gene. To illustrate this advantage of SAGE for functional genomics, we show that our data allow an unambiguous analysis of most of the individual genes belonging to 12 different ion transporter multigene families. These results indicate that, compared with EST-based tag to gene assignment, the use of the annotated genome sequence greatly improves gene identification in SAGE studies. However, more than 6,000 different tags remained with no gene match, suggesting that a significant proportion of transcripts present in the roots originate from yet unknown or wrongly annotated genes. The root transcriptome characterized in this study markedly differs from those obtained in other organs, and provides a unique resource for investigating the functional specificities of the root system. As an example of the use of SAGE for transcript profiling in Arabidopsis, we report here the identification of 270 genes differentially expressed between roots of plants grown either with NO3- or NH4NO3 as N source.

Arabidopsis↗

Theatre: A software tool for detailed comparative analysis and visualization of genomic sequence.

Theatre is a web-based computing system designed for the comparative analysis of genomic sequences, especially with respect to motifs likely to be involved in the regulation of gene expression. Theatre is an interface to commonly used sequence analysis tools and biological sequence databases to determine or predict the positions of coding regions, repetitive sequences and transcription factor binding sites in families of DNA sequences. The information is displayed in a manner that can be easily understood and can reveal patterns that might not otherwise have been noticed. In addition to web-based output, Theatre can produce publication quality colour hardcopies showing predicted features in aligned genomic sequences. A case study using the p53 promoter region of four mammalian species and two fish species is described. Unlike the mammalian sequences the promoter regions in fish have not been previously predicted or characterized and we report the differences in the p53 promoter region of four mammals and that predicted for two fish species. Theatre can be accessed at http://www.hgmp.mrc.ac.uk/Registered/Webapp/theatre/.

Animals↗

Genome sequence of the chemolithoautotrophic nitrite-oxidizing bacterium Nitrobacter winogradskyi Nb-255.

The alphaproteobacterium Nitrobacter winogradskyi (ATCC 25391) is a gram-negative facultative chemolithoautotroph capable of extracting energy from the oxidation of nitrite to nitrate. Sequencing and analysis of its genome revealed a single circular chromosome of 3,402,093 bp encoding 3,143 predicted proteins. There were extensive similarities to genes in two alphaproteobacteria, Bradyrhizobium japonicum USDA110 (1,300 genes) and Rhodopseudomonas palustris CGA009 CG (815 genes). Genes encoding pathways for known modes of chemolithotrophic and chemoorganotrophic growth were identified. Genes encoding multiple enzymes involved in anapleurotic reactions centered on C2 to C4 metabolism, including a glyoxylate bypass, were annotated. The inability of N. winogradskyi to grow on C6 molecules is consistent with the genome sequence, which lacks genes for complete Embden-Meyerhof and Entner-Doudoroff pathways, and active uptake of sugars. Two gene copies of the nitrite oxidoreductase, type I ribulose-1,5-bisphosphate carboxylase/oxygenase, cytochrome c oxidase, and gene homologs encoding an aerobic-type carbon monoxide dehydrogenase were present. Similarity of nitrite oxidoreductases to respiratory nitrate reductases was confirmed. Approximately 10% of the N. winogradskyi genome codes for genes involved in transport and secretion, including the presence of transporters for various organic-nitrogen molecules. The N. winogradskyi genome provides new insight into the phylogenetic identity and physiological capabilities of nitrite-oxidizing bacteria. The genome will serve as a model to study the cellular and molecular processes that control nitrite oxidation and its interaction with other nitrogen-cycling processes.

Bacterial Proteins↗

Complete genome sequence of the oral pathogenic Bacterium porphyromonas gingivalis strain W83.

The complete 2,343,479-bp genome sequence of the gram-negative, pathogenic oral bacterium Porphyromonas gingivalis strain W83, a major contributor to periodontal disease, was determined. Whole-genome comparative analysis with other available complete genome sequences confirms the close relationship between the Cytophaga-Flavobacteria-Bacteroides (CFB) phylum and the green-sulfur bacteria. Within the CFB phyla, the genomes most similar to that of P. gingivalis are those of Bacteroides thetaiotaomicron and B. fragilis. Outside of the CFB phyla the most similar genome to P. gingivalis is that of Chlorobium tepidum, supporting the previous phylogenetic studies that indicated that the Chlorobia and CFB phyla are related, albeit distantly. Genome analysis of strain W83 reveals a range of pathways and virulence determinants that relate to the novel biology of this oral pathogen. Among these determinants are at least six putative hemagglutinin-like genes and 36 previously unidentified peptidases. Genome analysis also reveals that P. gingivalis can metabolize a range of amino acids and generate a number of metabolic end products that are toxic to the human host or human gingival tissue and contribute to the development of periodontal disease.

Bacterial Proteins↗

In silico prediction of scaffold/matrix attachment regions in large genomic sequences.

Scaffold/matrix attachment regions (S/MARs) are essential regulatory DNA elements of eukaryotic cells. They are major determinants of locus control of gene expression and can shield gene expression from position effects. Experimental detection of S/MARs requires substantial effort and is not suitable for large-scale screening of genomic sequences. In silico prediction of S/MARs can provide a crucial first selection step to reduce the number of candidates. We used experimentally defined S/MAR sequences as the training set and generated a library of new S/MAR-associated, AT-rich patterns described as weight matrices. A new tool called SMARTest was developed that identifies potential S/MARs by performing a density analysis based on the S/MAR matrix library (http://www.genomatix.de/cgi-bin/smartest_pd/smartest.pl). S/MAR predictions were evaluated by using six genomic sequences from animal and plant for which S/MARs and non-S/MARs were experimentally mapped. SMARTest reached a sensitivity of 38% and a specificity of 68%. In contrast to previous algorithms, the SMARTest approach does not depend on the sequence context and is suitable to analyze long genomic sequences up to the size of whole chromosomes. To demonstrate the feasibility of large-scale S/MAR prediction, we analyzed the recently published chromosome 22 sequence and found 1198 S/MAR candidates.

Algorithms↗

Slc19a2: cloning and characterization of the murine thiamin transporter cDNA and genomic sequence, the orthologue of the human TRMA gene.

Recently, our group and others cloned the TRMA disease gene, SLC19A2, which encodes a thiamin transporter. Here, we report the cloning and characterization of the full-length cDNA and genomic sequences of mouse Slc19a2. The Slc19a2 cDNA contained a 1494-bp open-reading frame, and had 5'- and 3'-untranslated regions of 189 and 1857 bp, respectively. A putative GC-rich, TATA-less promoter was identified in genomic sequence directly upstream of the identified 5' end. The Slc19a2 gene spanned 16.3 kb and was organized into six exons, a gene structure conserved with the human orthologue. The predicted Slc19a2 protein, like SLC19A2, was predicted to have 12 transmembrane domains and shared a number of other conserved sequence motifs with the human orthologue, including one potential N-glycosylation site (N(63)) and several potential phosphorylation sites. Comparison of the Slc19a2 amino acid sequence with those of the other known SLC19A solute carriers highlighted interesting patterns of conservation and divergence in various domains, allowing insight into potential structure-function relationships. The identification of the mouse Slc19a2 cDNA and genomic sequences will facilitate the generation of an animal model of TRMA, permitting future studies of disease pathogenesis.

Amino Acid Sequence↗

Representation of cloned genomic sequences in two sequencing vectors: correlation of DNA sequence and subclone distribution.

Representation of subcloned Caenorhabditis elegans and human DNA sequences in both M13 and pUC sequencing vectors was determined in the context of large scale genomic sequencing. In many cases, regions of subclone under-representation correlated with the occurrence of repeat sequences, and in some cases the under-representation was orientation specific. Factors which affected subclone representation included the nature and complexity of the repeat sequence, as well as the length of the repeat region. In some but not all cases, notable differences between the M13 and pUC subclone distributions existed. However, in all regions lacking one type of subclone (either M13 or pUC), an alternate subclone was identified in at least one orientation. This suggests that complementary use of M13 and pUC subclones would provide the most comprehensive subclone coverage of a given genomic sequence.

Animals↗

Complete plastid genome sequences of Drimys, Liriodendron, and Piper: implications for the phylogenetic relationships of magnoliids.

BACKGROUND: The magnoliids with four orders, 19 families, and 8,500 species represent one of the largest clades of early diverging angiosperms. Although several recent angiosperm phylogenetic analyses supported the monophyly of magnoliids and suggested relationships among the orders, the limited number of genes examined resulted in only weak support, and these issues remain controversial. Furthermore, considerable incongruence resulted in phylogenetic reconstructions supporting three different sets of relationships among magnoliids and the two large angiosperm clades, monocots and eudicots. We sequenced the plastid genomes of three magnoliids, Drimys (Canellales), Liriodendron (Magnoliales), and Piper (Piperales), and used these data in combination with 32 other angiosperm plastid genomes to assess phylogenetic relationships among magnoliids and to examine patterns of variation of GC content. RESULTS: The Drimys, Liriodendron, and Piper plastid genomes are very similar in size at 160,604, 159,886 bp, and 160,624 bp, respectively. Gene content and order are nearly identical to many other unrearranged angiosperm plastid genomes, including Calycanthus, the other published magnoliid genome. Overall GC content ranges from 34-39%, and coding regions have a substantially higher GC content than non-coding regions. Among protein-coding genes, GC content varies by codon position with 1st codon > 2nd codon > 3rd codon, and it varies by functional group with photosynthetic genes having the highest percentage and NADH genes the lowest. Phylogenetic analyses using parsimony and likelihood methods and sequences of 61 protein-coding genes provided strong support for the monophyly of magnoliids and two strongly supported groups were identified, the Canellales/Piperales and the Laurales/Magnoliales. Strong support is reported for monocots and eudicots as sister clades with magnoliids diverging before the monocot-eudicot split. The trees also provided moderate or strong support for the position of Amborella as sister to a clade including all other angiosperms. CONCLUSION: Evolutionary comparisons of three new magnoliid plastid genome sequences, combined with other published angiosperm genomes, confirm that GC content is unevenly distributed across the genome by location, codon position, and functional group. Furthermore, phylogenetic analyses provide the strongest support so far for the hypothesis that the magnoliids are sister to a large clade that includes both monocots and eudicots.

Base Composition↗

A genome sequence survey of the mollicute corn stunt spiroplasma Spiroplasma kunkelii.

The mollicute corn stunt spiroplasma (Spiroplasma kunkelii) is a leafhopper-transmitted pathogen of maize. Sequencing of the approximately 1.6-Mb genome of S. kunkelii was initiated to aid understanding the genetic basis of spiroplasma interactions with their plant and leafhopper hosts. In total, 144712 nucleotides of non-redundant, high-quality S. kunkelii genome sequence were obtained. Sequence tags were searched against the Mycoplasmataceae and Bacillus/Clostridium databases. Results showed that, in addition to spiroplasma phage SpV1 DNA insertions, spiroplasma genomes harbor more purine and amino acid biosynthesis, transcription regulation, cell envelope and DNA transport/binding genes than Mycoplasmataceae genomes. This investigation demonstrates that survey sequencing is an efficient procedure for gene discovery and genome characterization. The results of the S. kunkelii sequencing project are available at the Spiroplasma WebPage at http://www.oardc.ohio-state.edu/spiroplasma/genome.htm.

DNA Replication↗

Large-scale simulation of coverage and error rate tradeoffs for cancer detection in cell-free DNA whole-genome sequencing.

MOTIVATION: Cell-free DNA (cfDNA) whole-genome sequencing (WGS) is a promising approach for detecting cancer recurrence. It enables cancer detection by identifying all tumor-derived cfDNA (ctDNA) molecules carrying somatic single nucleotide variants (sSNVs). While ideally, a sequencing platform should be highly accurate for reliable ctDNA detection, in reality, all sequencing platforms introduce sequencing errors that generate false positives indistinguishable from true SNVs. Understanding how sequencing parameters influence ctDNA detection sensitivity at low tumor fractions (TFs) in cfDNA samples is essential for guiding sequencing strategies in clinical contexts. To model cfDNA sequencing for tumor detection, which contains asymmetric noise and multiple interacting parameters, analytical modeling is intractable, motivating large-scale parallelized simulation. RESULTS: We developed a simulation framework to generate in silico cfDNA data across 10 cancer types. In total, 480 million cfDNA samples were simulated from tumor WGS profiles. Overall, the lowest detectable TF differs substantially between cancer types under identical sequencing conditions due to variations in mutational load. For cancers with high mutational load, 3&#xd7; coverage with low-error techniques reliably detects TFs below 0.1%. In contrast, cancers with low mutational load require at least six-fold higher coverage to achieve comparable detection thresholds. Increasing sequencing quality scores from Q30 to Q55 at 30&#xd7; coverage further enhances sensitivity, enabling detection of TFs as low as 1&#x2009;&#xd7;&#x2009;10-5. This study provides a comprehensive framework for optimizing sequencing parameters, offering valuable guidance for tailoring future technology development for specific cancer types and clinical applications. AVAILABILITY AND IMPLEMENTATION: The code is publicly available at https://github.com/UMCUGenetics/cfdetect/tree/main.

Whole Genome Sequencing↗