Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

Short-insert libraries as a method of problem solving in genome sequencing.

As the Human Genome Project moves into its sequencing phase, a serious problem has arisen. The same problem has been increasingly vexing in the closing phase of the Caenorhabditis elegans project. The difficulty lies in sequencing efficiently through certain regions in which the templates (DNA substrates for the sequencing process) form complex folded secondary structures that are inaccessible to the enzymes. The solution, however, is simply to break them up. Specifically, the offending fragments are sonicated heavily and recloned, as much smaller fragments, into pUC vector. The sequences obtained from the resulting library can subsequently be assembled, free from the effects of secondary structure, to produce high-quality, complete sequence. Because of the success and simplicity of this procedure, we have begun to use it for the sequencing of all regions in which standard primer walking has been at all difficult.

Animals↗

Genomic sequencing reveals absence of DNA methylation in the major late promoter of adenovirus type 2 DNA in the virion and in productively infected cells.

By using methylation-sensitive restriction endonucleases, we have previously provided evidence that adenovirus type 2 (Ad2) virion DNA or free intranuclear Ad2 DNA in productively infected hamster or human cells is not methylated. We have now chosen a different experimental approach and have investigated the major late promoter (MLP) sequence of Ad2 DNA for the presence of 5-methyldeoxycytidine (5-mC) residues with the genomic sequencing technique. This study has been prompted by the finding that the MLP of Ad2 DNA can be inactivated by sequence-specific methylation in experiments in which a MLP-chloramphenicol acetyltransferase construct has been transcribed in a cell-free system from HeLa cell nuclear extracts. Virion Ad2 DNA and Ad2 DNA isolated from productively infected human or hamster cells between 1 and 48 h post-infection (p.i.) have now been analyzed. There is no evidence for the presence of 5-mC in the cytidine positions in the MLP of any of these Ad2 preparations. We conclude that DNA methylation does not seem to play a role in the early-late control of this viral promoter. The sensitivity of the genomic sequencing technique does not permit us to exclude the unlikely presence of 5-mC in a few Ad2 DNA molecules.

Adenoviruses, Human↗

Genomic sequence analysis of Epstein-Barr virus strain GD1 from a nasopharyngeal carcinoma patient.

To date, the only entire Epstein-Barr virus (EBV) genomic sequence available in the database is the prototype B95.8, which was derived from an individual with infectious mononucleosis. A causative link between EBV and nasopharyngeal carcinoma (NPC), a disease with a distinctly high incidence in southern China, has been widely investigated. However, no full-length analysis of any substrain of EBV from this area has been reported. In this study, we analyzed the entire genomic sequence of an EBV strain from a patient with NPC in Guangdong, China. This EBV strain was termed GD1 (Guangdong strain 1), and the full-length sequence of GD1 was submitted to the GenBank database. The assigned accession number is AY961628. The entire GD1 sequence is 171,656 bp in length, with 59.5% G+C content and 40.5% A+T content. We detected many sequence variations in GD1 compared to prototypical strain B95.8, including 43 deletion sites, 44 insertion sites, and 1,413 point mutations. Furthermore, we evaluated the frequency of some of these GD1 mutations in Cantonese NPC patients and found them to be highly prevalent. These findings suggest that GD1 is highly representative of the EBV strains isolated from NPC patients in Guangdong, China, an area with the highest incidence of NPC in the world. Furthermore, these findings provide the second full-length sequence analysis of any EBV strain as well as the first full-length sequence analysis of an NPC-derived EBV strain.

Adult↗

The genome sequence of Sphagnum tenellum (Brid.) Bory, 1819 (Sphagnales: Sphagnaceae).

We present a genome assembly of Sphagnum tenellum (soft bog-moss; Streptophyta; Sphagnopsida; Sphagnales; Sphagnaceae). The genome sequence has a total length of 386.16 megabases. Most of the assembly (98.04%) is scaffolded into 21 chromosomal pseudomolecules. The mitochondrial sequence has a length of 141.31 kilobases and the plastid genome assembly has a length of 140.16 kilobases. This assembly was generated as part of the Darwin Tree of Life project, which produces reference genomes for eukaryotic species found in Britain and Ireland.

Sphagnales↗

IsoFinder: computational prediction of isochores in genome sequences.

Isochores are long genome segments homogeneous in G+C. Here, we describe an algorithm (IsoFinder) running on the web (http://bioinfo2.ugr.es/IsoF/isofinder.html) able to predict isochores at the sequence level. We move a sliding pointer from left to right along the DNA sequence. At each position of the pointer, we compute the mean G+C values to the left and to the right of the pointer. We then determine the position of the pointer for which the difference between left and right mean values (as measured by the t-statistic) reaches its maximum. Next, we determine the statistical significance of this potential cutting point, after filtering out short-scale heterogeneities below 3 kb by applying a coarse-graining technique. Finally, the program checks whether this significance exceeds a probability threshold. If so, the sequence is cut at this point into two subsequences; otherwise, the sequence remains undivided. The procedure continues recursively for each of the two resulting subsequences created by each cut. This leads to the decomposition of a chromosome sequence into long homogeneous genome regions (LHGRs) with well-defined mean G+C contents, each significantly different from the G+C contents of the adjacent LHGRs. Most LHGRs can be identified with Bernardi's isochores, given their correlation with biological features such as gene density, SINE and LINE (short, long interspersed repetitive elements) densities, recombination rate or single nucleotide polymorphism variability. The resulting isochore maps are available at our web site (http://bioinfo2.ugr.es/isochores/), and also at the UCSC Genome Browser (http://genome.cse.ucsc.edu/).

Algorithms↗

Prediction of transcriptional regulatory sites in the complete genome sequence of Escherichia coli K-12.

MOTIVATION: As one of the best-characterized free-living organisms, Escherichia coli and its recently completed genomic sequence offer a special opportunity to exploit systematically the variety of regulatory data available in the literature in order to make a comprehensive set of regulatory predictions in the whole genome. RESULTS: The complete genome sequence of E.coli was analyzed for the binding of transcriptional regulators upstream of coding sequences. The biological information contained in RegulonDB (Huerta, A.M. et al., Nucleic Acids Res.,26,55-60, 1998) for 56 different transcriptional proteins was the support to implement a stringent strategy combining string search and weight matrices. We estimate that our search included representatives of 15-25% of the total number of regulatory binding proteins in E.coli. This search was performed on the set of 4288 putative regulatory regions, each 450 bp long. Within the regions with predicted sites, 89% are regulated by one protein and 81% involve only one site. These numbers are reasonably consistent with the distribution of experimental regulatory sites. Regulatory sites are found in 603 regions corresponding to 16% of operon regions and 10% of intra-operonic regions. Additional evidence gives stronger support to some of these predictions, including the position of the site, biological consistency with the function of the downstream gene, as well as genetic evidence for the regulatory interaction. The predictions described here were incorporated into the map presented in the paper describing the complete E.coli genome (Blattner,F.R. et al., Science, 277, 1453-1461, 1997). AVAILABILITY: The complete set of predictions in GenBank format is available at the url: http://www. cifn.unam.mx/Computational_Biology/E.coli-predictions CONTACT: ecoli-reg@cifn.unam.mx, collado@cifn.unam.mx

Bacterial Proteins↗

The genome sequence of Hyocomium armoricum (Brid.) Wijk & Margad. (Hypnales: Hypnaceae).

We present a genome assembly of Hyocomium armoricum (Flagellate Feather-moss; Streptophyta; Bryopsida; Hypnales; Hypnaceae). The genome sequence has a total length of 333.60 megabases. Most of the assembly (99.59%) is scaffolded into 11 chromosomal pseudomolecules. The mitochondrial sequence has a length of 104.43 kilobases and the plastid genome assembly has a length of 123.84 kilobases. This assembly was generated as part of the Darwin Tree of Life project, which produces reference genomes for eukaryotic species found in Britain and Ireland.

Flagellate Feather-moss↗

Genome sequence completed of Alcanivorax borkumensis, a hydrocarbon-degrading bacterium that plays a global role in oil removal from marine systems.

In this paper, we provide background to the genome sequencing project of Alcanivorax borkumensis, which is a marine bacterium that uses exclusively petroleum oil hydrocarbons as sources of carbon and energy (therefore designated "hydrocarbonoclastic"). It is found in low numbers in all oceans of the world and in high numbers in oil-contaminated waters. Its ubiquity and unusual physiology suggest it is globally important in the removal of hydrocarbons from polluted marine systems. A functional genomics analysis of Alcanivorax borkumensis strain SK2 was recently initiated, and its genome sequence has just been completed. Annotation of the genome, metabolome modelling, and functional genomics, will soon reveal important insights into the genomic basis of the properties and physiology of this fascinating and globally important bacterium.

Biodegradation, Environmental↗

Whole-genome sequencing in 333,100 individuals reveals rare non-coding single variant and aggregate associations with height.

The role of rare non-coding variation in complex human phenotypes is still largely unknown. To elucidate the impact of rare variants in regulatory elements, we performed a whole-genome sequencing association analysis for height using 333,100 individuals from three datasets: UK Biobank (N&#x2009;=&#x2009;200,003), TOPMed (N&#x2009;=&#x2009;87,652) and All of Us (N&#x2009;=&#x2009;45,445). We performed rare (&#x2009;<&#x2009;0.1% minor-allele-frequency) single-variant and aggregate testing of non-coding variants in regulatory regions based on proximal-regulatory, intergenic-regulatory and deep-intronic annotation. We observed 29 independent variants associated with height at P&#x2009;<&#x2009;after conditioning on previously reported variants, with effect sizes ranging from -7cm to +4.7&#x2009;cm. We also identified and replicated non-coding aggregate-based associations proximal to HMGA1 containing variants associated with a 5&#x2009;cm taller height and of highly-conserved variants in MIR497HG on chromosome 17. We have developed an approach for identifying non-coding rare variants in regulatory regions with large effects from whole-genome sequencing data associated with complex traits.

Humans↗

The complete genome sequence of severe acute respiratory syndrome coronavirus strain HKU-39849 (HK-39).

The complete genomic nucleotide sequence (29.7kb) of a Hong Kong severe acute respiratory syndrome (SARS) coronavirus (SARS-CoV) strain HK-39 is determined. Phylogenetic analysis of the genomic sequence reveals it to be a distinct member of the Coronaviridae family. 5' RACE assay confirms the presence of at least six subgenomic transcripts all containing the predicted intergenic sequences. Five open reading frames (ORFs), namely ORF1a, 1b, S, M, and N, are found to be homologues to other CoV members, and three more unknown ORFs (X1, X2, and X3) are unparalleled in all other known CoV species. Optimal alignment and computer analysis of the homologous ORFs has predicted the characteristic structural and functional domains on the putative genes. The overall nucleotides conservation of the homologous ORFs is low (<5%) compared with other known CoVs, implying that HK-39 is a newly emergent SARS-CoV phylogenetically distant from other known members. SimPlot analysis supports this finding, and also suggests that this novel virus is not a product of a recent recombinant from any of the known characterized CoVs. Together, these results confirm that HK-39 is a novel and distinct member of the Coronaviridae family, with unknown origin. The completion of the genomic sequence of the virus will assist in tracing its origin.

3' Untranslated Regions↗

Highly specific localization of promoter regions in large genomic sequences by PromoterInspector: a novel context analysis approach.

We present a new algorithm called PromoterInspector to locate eukaryotic polymase II promoter regions in large genomic sequences with a high degree of specificity. PromoterInspector focuses on the genetic context of promoters, rather than their exact location. Application of PromoterInspector can serve as a crucial pre-processing step for other methods to locate exactly, or to analyze promoters. PromoterInspector does not depend on heuristics, because it is purely based on libraries of IUPAC words extracted from training sequences by an unsupervised learning approach. We compared PromoterInspector to in silico promoter prediction tools using the sequences from the review by J.W. Fickett. PromoterInspector compared favourably on Fickett's evaluation scheme. A true positive to false positive ratio of 2.3 was obtained, surpassing the best ratio of 0.6, reported for TSSG. The application of our method to several large genomic sequences of over 1.3 million base-pairs in total resulted in even more specific predictions. The coverage of annotated promoters was comparable to other in silico promoter prediction methods, while the true positive predictions increased by up to 100% of total matches. PromoterInspector scans 100 kb in less than one minute on a workstation, and thus is especially applicable for large genome analysis. The method is available at http://genomatix.gsf. de/cgi-bin/promoterinspector/promoterinspector.pl.

3' Untranslated Regions↗

Analysis of the genomic sequence of a human metapneumovirus.

We recently described the isolation of a novel paramyxovirus from children with respiratory tract disease in The Netherlands. Based on biological properties and limited sequence information the virus was provisionally classified as the first nonavian member of the Metapneumovirus genus and named human metapneumovirus (hMPV). This report describes the analysis of the sequences of all hMPV open reading frames (ORFs) and intergenic sequences as well as partial sequences of the genomic termini. The overall percentage of amino acid sequence identity between APV and hMPV N, P, M, F, M2-1, M2-2, and L ORFs was 56 to 88%. Some nucleotide sequence identity was also found between the noncoding regions of the APV and hMPV genomes. Although no discernible amino acid sequence identity was found between two of the ORFs of hMPV and ORFs of other paramyxoviruses, the amino acid content, hydrophilicity profiles, and location of these ORFs in the viral genome suggest that they represent SH and G proteins. The high percentage of sequence identity between APV and hMPV, their similar genomic organization (3'-N-P-M-F-M2-SH-G-L-5'), and phylogenetic analyses provide evidence for the proposed classification of hMPV as the first mammalian metapneumovirus.

Amino Acid Sequence↗

The genome sequence of a caddisfly, Limnephilus auricula (Curtis, 1834).

We present a genome assembly from an individual female Limnephilus auricula (a caddisfly; Arthropoda; Insecta; Trichoptera; Limnephilidae). The genome sequence is 971.3 megabases in span. Most of the assembly is scaffolded into 30 chromosomal pseudomolecules, including the Z sex chromosome. The mitochondrial genome has also been assembled and is 18.29 kilobases in length.

Limnephilus auricula↗

The complete genome sequence of an El Amar isolate of plum pox virus (PPV) and its phylogenetic relationship to other PPV strains.

The genomic sequence of an El Amar isolate of plum pox virus (PPV) from Egypt was determined by sequencing overlapping cDNA fragments. This is the first complete sequence of a member of the El Amar (EA) strain of PPV. The genome consists of 9791 nt, excluding a poly(A) tail at the 3' terminus. The complete nt sequence of PPV EA is 79-80%, 80%, 77%, and 77% homologous with isolates of strains D/M, Rec (BOR3), C, and W, respectively. The polyprotein identity ranged from 87-91%. Phylogenetic analysis using the complete genome sequence of PPV EA confirmed its strain status. No significant recombination signals were identified using PhylPro and SimPlot scans of the PPV EA sequence, however an interesting recombination signal was identified in the P1/HC-Pro region of PPV W3174.

Base Sequence↗

The genome sequence of Caenorhabditis briggsae: a platform for comparative genomics.

The soil nematodes Caenorhabditis briggsae and Caenorhabditis elegans diverged from a common ancestor roughly 100 million years ago and yet are almost indistinguishable by eye. They have the same chromosome number and genome sizes, and they occupy the same ecological niche. To explore the basis for this striking conservation of structure and function, we have sequenced the C. briggsae genome to a high-quality draft stage and compared it to the finished C. elegans sequence. We predict approximately 19,500 protein-coding genes in the C. briggsae genome, roughly the same as in C. elegans. Of these, 12,200 have clear C. elegans orthologs, a further 6,500 have one or more clearly detectable C. elegans homologs, and approximately 800 C. briggsae genes have no detectable matches in C. elegans. Almost all of the noncoding RNAs (ncRNAs) known are shared between the two species. The two genomes exhibit extensive colinearity, and the rate of divergence appears to be higher in the chromosomal arms than in the centers. Operons, a distinctive feature of C. elegans, are highly conserved in C. briggsae, with the arrangement of genes being preserved in 96% of cases. The difference in size between the C. briggsae (estimated at approximately 104 Mbp) and C. elegans (100.3 Mbp) genomes is almost entirely due to repetitive sequence, which accounts for 22.4% of the C. briggsae genome in contrast to 16.5% of the C. elegans genome. Few, if any, repeat families are shared, suggesting that most were acquired after the two species diverged or are undergoing rapid evolution. Coclustering the C. elegans and C. briggsae proteins reveals 2,169 protein families of two or more members. Most of these are shared between the two species, but some appear to be expanding or contracting, and there seem to be as many as several hundred novel C. briggsae gene families. The C. briggsae draft sequence will greatly improve the annotation of the C. elegans genome. Based on similarity to C. briggsae, we found strong evidence for 1,300 new C. elegans genes. In addition, comparisons of the two genomes will help to understand the evolutionary forces that mold nematode genomes.

Animals↗

The genome sequence of the Scarce Umber, Agriopis aurantiaria (H&#xfc;bner, 1799).

We present a genome assembly from an individual male Agriopis aurantiaria (the Scarce Umber; Arthropoda; Insecta; Lepidoptera; Geometridae). The genome sequence is 485.4 megabases in span. The whole assembly is scaffolded into 30 chromosomal pseudomolecules, including the Z sex chromosome. The mitochondrial genome has also been assembled and is 15.44 kilobases in length. Gene annotation of this assembly on Ensembl identified 16,963 protein coding genes.

Agriopis aurantiaria↗