Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Streamlining large-scale genomic data management: Insights from the UK Biobank whole-genome sequencing data.

Biobank-scale whole-genome sequencing (WGS) studies are increasingly pivotal in unraveling the genetic bases of diverse health outcomes. However, managing and analyzing these datasets' sheer volume and complexity presents significant challenges. We highlight the annotated genomic data structure (aGDS) format, substantially reducing the WGS data file size while enabling seamless integration of genomic and functional information for comprehensive WGS analyses. The aGDS format yielded 23 chromosome-specific files for the UK Biobank 500k WGS dataset, occupying only 1.10 tebibytes of storage. We develop the vcf2agds toolkit that streamlines the conversion of WGS data from VCF to aGDS format. Additionally, the STAARpipeline equipped with the aGDS files enabled scalable, comprehensive, and functionally informed WGS analysis, facilitating the detection of common and rare coding and noncoding phenotype-genotype associations. Overall, the vcf2agds toolkit and STAARpipeline provide a streamlined solution that facilitates efficient data management and analysis of biobank-scale WGS data across hundreds of thousands of samples.

Humans↗

Comparisons of eukaryotic genomic sequences.

A method for assessing genomic similarity based on relative abundances of short oligonucleotides in large DNA samples is introduced. The method requires neither homologous sequences nor prior sequence alignments. The analysis centers on (i) dinucleotide (and tri- and tetra-) relative abundance extremes in genomic sequences, (ii) distances between sequences based on all dinucleotide relative abundance values, and (iii) a multidimensional partial ordering protocol. The emphasis in this paper is on assessments of general relatedness of genomes as distinguished from phylogenetic reconstructions. Our methods demonstrate that the relative abundance distances almost always differ more for genomic interspecific sequence comparisons than for genomic intraspecific sequence comparisons, indicating congruence over different genome sequence samples. The genomic comparisons are generally concordant with accepted phylogenies among vertebrate and among fungal species sequences. Several unexpected relationships between the major groups of metazoa, fungal, and protist DNA emerge, including the following. (i) Schizosaccharomyces pombe and Saccharomyces cerevisiae in dinucleotide relative abundance distances are as similar to each other as human is to bovine. (ii) S. cerevisiae, although substantially far from, is significantly closer to the vertebrates than are the invertebrates (Drosophila melanogaster, Bombyx mori, and Caenorhabditis elegans). This phenomenon may suggest variable evolutionary rates during the metazoan radiations and slower changes in the fungal divergences, and/or a polyphyletic origin of metazoa. (iii) The genomic sequences of D. melanogaster and Trypanosoma brucei are strikingly similar. This DNA similarity might be explained by some molecular adaptation of the parasite to its dipteran (tsetse fly) host, a host-parasite gene transfer hypothesis. Robustness of the methods may be due to a genomic signature of dinucleotide relative abundance values reflecting DNA structures related to dinucleotide stacking energies, constraints of DNA curvature, and mechanisms attendant to replication, repair, and recombination.

Animals↗

Updating of transposable element annotations from large wheat genomic sequences reveals diverse activities and gene associations.

Triticeae species (including wheat, barley and rye) have huge and complex genomes due to polyploidization and a high content of transposable elements (TEs). TEs are known to play a major role in the structure and evolutionary dynamics of Triticeae genomes. During the last 5 years, substantial stretches of contiguous genomic sequence from various species of Triticeae have been generated, making it necessary to update and standardize TE annotations and nomenclature. In this study we propose standard procedures for these tasks, based on structure, nucleic acid and protein sequence homologies. We report statistical analyses of TE composition and distribution in large blocks of genomic sequences from wheat and barley. Altogether, 3.8 Mb of wheat sequence available in the databases was analyzed or re-analyzed, and compared with 1.3 Mb of re-annotated genomic sequences from barley. The wheat sequences were relatively gene-rich (one gene per 23.9 kb), although wheat gene-derived sequences represented only 7.8% (159 elements) of the total, while the remainder mainly comprised coding sequences found in TEs (54.7%, 751 elements). Class I elements [mainly long terminal repeat (LTR) retrotransposons] accounted for the major proportion of TEs, in terms of sequence length as well as element number (83.6% and 498, respectively). In addition, we show that the gene-rich sequences of wheat genome A seem to have a higher TE content than those of genomes B and D, or of barley gene-rich sequences. Moreover, among the various TE groups, MITEs were most often associated with genes: 43.1% of MITEs fell into this category. Finally, the TRIM and copia elements were shown to be the most active TEs in the wheat genome. The implications of these results for the evolution of diploid and polyploid wheat species are discussed.

DNA Transposable Elements↗

Multiple alignment of genomic sequences using CHAOS, DIALIGN and ABC.

Comparative analysis of genomic sequences is a powerful approach to discover functional sites in these sequences. Herein, we present a WWW-based software system for multiple alignment of genomic sequences. We use the local alignment tool CHAOS to rapidly identify chains of pairwise similarities. These similarities are used as anchor points to speed up the DIALIGN multiple-alignment program. Finally, the visualization tool ABC is used for interactive graphical representation of the resulting multiple alignments. Our software is available at Göttingen Bioinformatics Compute Server (GOBICS) at http://dialign.gobics.de/chaos-dialign-submission.

Computer Graphics↗

Genomic sequencing and methylation analysis by ligation mediated PCR.

Genomic sequencing permits studies of in vivo DNA methylation and protein-DNA interactions, but its use has been limited because of the complexity of the mammalian genome. A newly developed genomic sequencing procedure in which a ligation mediated polymerase chain reaction (PCR) is used generates high quality, reproducible sequence ladders starting with only 1 microgram of uncloned mammalian DNA per reaction. Different sequence ladders can be created simultaneously by inclusion of multiple primers and visualized separately by rehybridization. Relatively little radioactivity is needed for hybridization and exposure times are short. Methylation patterns in genomic DNA are readily detectable; for example, 17 CpG dinucleotides in the 5' region of human X-linked PGK-1 (phosphoglycerate kinase 1) were found to be methylated on an inactive human X chromosome, but unmethylated on an active X chromosome.

5-Methylcytosine↗

Of mice and genome sequence.

Availability of the mouse genome sequence will have a major impact on the study of vertebrate evolution, mammalian biology, and animal models of human disease. Resources to explore genome biology in mice will maximize the effect of this watershed event.

Animals↗

Identifying tagged transposon insertion sites in yeast by direct genomic sequencing.

Tagged transposons are powerful tools for large-scale studies of gene expression, protein localization, and gene disruption in Saccharomyces cerevisiae. The current techniques used to identify transposon insertion sites in the yeast genome require a DNA amplification step that can be time-consuming and problematic. We show that the DNA amplification step can be bypassed. Insertion sites can be identified rapidly and reliably by direct genomic sequencing using a transposon-specific primer, BigDye-labelled terminators, and an automated sequencer. Direct genomic sequencing can also save time on the genetic analysis phase of transposon-based projects.

Base Sequence↗

Complete genomic sequence and phylogenetic relatedness of hepatitis B virus isolates from Iran.

Hepatitis B virus (HBV) is one of the main etiological agents of acute and chronic liver disease that is still a major public health problem in the world. Numerous HBV isolates have grouped into eight genotypes, A to H, based on the complete genome sequence. To date, no study has been carried out on the complete HBV genome sequence in Iran. The objective of this study was to investigate the complete genome sequence organization and phylogenetic analysis of the five HBV strains, which obtained from Iranian chronic infected patients. Results showed that Iranian strains were closely related to each other, with 97-100% nucleotide similarity. Phylogenetic analysis based on the complete genome sequences and the precore/core gene sequences revealed that all strains were of genotype D, sub-genotype D1 with bootstrap value 100 and 99%, respectively. The S gene encoded Arg122, Pro127, and Lys160 corresponding to subtype ayw2. Iranian HBV isolates had closely related with Turkish HBV strains. All strains had a nucleotide length of 3,182 base pair (bp) except IR-P4 strain, with a 3,185 bp in length and with a unique Phe89 insertion in the X gene. The intragenotypic divergence of the complete genome sequence of Iranian strains was 1.8% and the intergenotypic in genotype D was 3.8% and with the other genotypes was 7.9-15.4%. In conclusion, this study revealed that the HBV genotype D, sub-genotype D1, subtype ayw2 dominates in the Iranian infected patients. A single Phe89 insertion in the X gene of the one Iranian strain with an unforeseen length of 3185 bp was identified.

Adult↗

Completion of the porcine epidemic diarrhoea coronavirus (PEDV) genome sequence.

The sequence of the replicase gene of porcine epidemic diarrhoea virus (PEDV) has been determined. This completes the sequence of the entire genome of strain CV777, which was found to be 28,033 nucleotides (nt) in length (excluding the poly A-tail). A cloning strategy, which involves primers based on conserved regions in the predicted ORF1 products from other coronaviruses whose genome sequence has been determined, was used to amplify the equivalent, but as yet unknown, sequence of PEDV. Primary sequences derived from these products were used to design additional primers resulting in the amplification and sequencing of the entire ORF1 of PEDV. Analysis of the nucleotide sequences revealed a small open reading frame (ORF) located near the 5' end (no 99-137), and two large, slightly overlapping ORFs, ORF1a (nt 297-12650) and ORF1b (nt 12605-20641). The ORF1a and ORF1b sequences overlapped at a potential ribosomal frame shift site. The amino acid sequence analysis suggested the presence of several functional motifs within the putative ORF1 protein. By analogy to other coronavirus replicase gene products, three protease and one growth factor-like motif were seen in ORF1a, and one polymerase domain, one metal ion-binding domain, and one helicase motif could be assigned within ORF1b. Comparative amino acid sequence alignments revealed that PEDV is most closely related to human coronavirus (HCoV)-229E and transmissible gastroenteritis virus (TGEV) and less related to murine hepatitis virus (MHV) and infectious bronchitis virus (IBV). These results thus confirm and extend the findings from sequence analysis of the structural genes of PEDV.

Amino Acid Sequence↗

Complete genomic sequence of the fish rhabdovirus infectious haematopoietic necrosis virus.

The complete nucleotide sequence of the genome of the fish rhabdovirus infectious haematopoietic necrosis virus (IHNV) has been determined after cDNA cloning of the viral genomic RNA. Sequence analysis showed the presence of six open reading frames encoding the nucleoprotein N, the matrix proteins M1 and M2, the glycoprotein G, a so-called non-structural protein NV, and the RNA polymerase L. The genome organization is 3'N-M1-M2-G-NV-L 5'. The extreme 5' and 3' ends of the genome were sequenced after RNA ligation or RACE. Prokaryotic expression products of the open reading frames predicted to encode the matrix proteins M1 and M2, the glycoprotein G and the NV protein reacted with rabbit anti-IHNV serum thereby confirming their identity. This is the first complete nucleotide sequence of a fish rhabdovirus. Knowledge of the complete sequence is an essential prerequisite for future manipulation of the genome and also serves to provide gene- and protein specific reagents for use in further examination of the replication of the fish rhabdoviruses.

Amino Acid Sequence↗

Human genomic sequences corresponding to murine CD3 eta-related transcripts: lack of conservation or expression of homologous human products.

We have cloned and sequenced human genomic DNA homologous to exons 9 and 10 of the CD3 zeta/eta/theta locus. Although there are open reading frames within the human sequences corresponding to the translated portions of murine exons 9 and 10, we find no evidence of conservation of the encoded polypeptide product. Furthermore, using oligonucleotides derived from these homologous sequences, we are unable to detect human CD3 eta- or CD3 theta-like transcripts by polymerase chain reaction amplification of reverse-transcribed RNA from a variety of human lymphoid tissues. Despite the absence of evidence for conservation of human CD3 eta and CD3 theta, there is a surprising degree of similarity between human and murine nucleotide sequences, not only for exons 9 and 10 (78% and 70%, respectively), but also for the 9/10 intron (71%). A possible mechanism for this conservation is discussed.

Animals↗

Characterization of the complete genomic sequence of genotype II hepatitis A virus (CF53/Berne isolate).

The complete genomic sequence of hepatitis A virus (HAV) CF53/Berne strain was determined. Pairwise comparison with other complete HAV genomic sequences demonstrated that the CF53/Berne isolate is most closely related to the single genotype VII strain, SLF88. This close relationship was confirmed by phylogenetic analyses of different genomic regions, and was most pronounced within the capsid region. These data indicated that CF53/Berne and SLF88 isolates are related more closely to each other than are subtypes IA and IB. A histogram of the genetic differences between HAV strains revealed four separate peaks. The distance values for CF53/Berne and SLF88 isolates fell within the peak that contained strains of the same subtype, showing that they should be subtypes within a single genotype. The complete genomic data indicated that genotypes II and VII should be considered a single genotype, based upon the complete VP1 sequence, and it is proposed that the CF53/Berne isolate be classified as genotype IIA and strain SLF88 as genotype IIB. The CF53/Berne isolate is cell-adapted, and therefore its sequence was compared to that of two other strains adapted to cell culture, HM-175/7 grown in MK-5 and GBM grown in FRhK-4 cells. Mutations found at nucleotides 3889, 4087 and 4222 that were associated with HAV attenuation and cell adaptation in HM175/7 and GMB strains were not present in the CF53/Berne strain. Deletions found in the 5'UTR and P3A regions of the CF53/Berne isolate that are common to cell-adapted HAV isolates were identified, however.

Base Sequence↗

Aligning multiple genomic sequences with the threaded blockset aligner.

We define a "threaded blockset," which is a novel generalization of the classic notion of a multiple alignment. A new computer program called TBA (for "threaded blockset aligner") builds a threaded blockset under the assumption that all matching segments occur in the same order and orientation in the given sequences; inversions and duplications are not addressed. TBA is designed to be appropriate for aligning many, but by no means all, megabase-sized regions of multiple mammalian genomes. The output of TBA can be projected onto any genome chosen as a reference, thus guaranteeing that different projections present consistent predictions of which genomic positions are orthologous. This capability is illustrated using a new visualization tool to view TBA-generated alignments of vertebrate Hox clusters from both the mammalian and fish perspectives. Experimental evaluation of alignment quality, using a program that simulates evolutionary change in genomic sequences, indicates that TBA is more accurate than earlier programs. To perform the dynamic-programming alignment step, TBA runs a stand-alone program called MULTIZ, which can be used to align highly rearranged or incompletely sequenced genomes. We describe our use of MULTIZ to produce the whole-genome multiple alignments at the Santa Cruz Genome Browser.

Animals↗

The mitochondrial genome of the mosquito Anopheles gambiae: DNA sequence, genome organization, and comparisons with mitochondrial sequences of other insects.

The entire 15,363 bp mitochondrial genome was cloned and sequenced from the mosquito Anopheles gambiae. With respect to the protein-coding genes, rRNA genes and the control region, the gene order was identical to that reported for other insects. There were significant differences, however, in the position and orientation of specific tRNA loci. The overall nucleotide composition was heavily biased towards adenine and thymine, which accounted for 77.6% of all nucleotides. Comparisons were made with the mitochondrial genomes of other insects on the basis genome size and organization, DNA and putative amino acid sequence data, nucleotide substitutions, codon usage and bias, and patterns of AT enrichment.

Amino Acid Sequence↗

Risk assessment prediction from genome sequences: promises and dreams.

The application of bacterial genomics opens new avenues of research on foodborne pathogens. Foodborne pathogens must be able to colonize their hosts and survive transmission from host to host. Different groups of genes are involved in the processes of survival, colonization, and virulence, and such genes are potential targets for risk assessment and intervention strategies. Filtering from genome sequences the genes relevant to these processes is a major challenge, and although many tools are already available for analyses, this type of data mining is just beginning. For the simplest application, gene comparison, it is important to know how gene function, for instance in virulence, is being defined and tested. In other genomic applications, reserachers look for specific properties or characteristics of (virulence) genes to identify novel gene candidates. Each approach has pitfalls, and gene candidates must be tested in the lab to confirm their function. Models for colonization and virulence are available for most although not all pathogens. Models for survival and stress responses are needed to increase the utilization of genomic approaches to risk assessment. Here, I discuss how genome sequences are likely to help in microbial risk assessment of foodborne pathogens and how dreams may become promises.

Animals↗

Equity in genome sequencing for rare disease diagnosis: a cross-sectional analysis of data from the UK 100,000 Genomes Project.

BACKGROUND: Genome sequencing has improved rare disease diagnosis and is now part of routine clinical care in the National Health Service in England. Automated prioritisation pipelines narrow millions of variants per patient to a small subset for clinical review, a process that relies on allele frequency resources that do not fully represent human genetic diversity. We assessed ancestry-related differences in variant prioritisation and diagnostic outcomes in patients from the UK 100,000 Genomes Project. METHODS: We analysed 29,405 rare disease probands with genome sequencing and linked clinical outcomes data. We used multivariable regression to assess ancestry-related differences in the number of variants prioritised for clinical review, the proportion of prioritised variants that were recorded as diagnostic, and diagnostic yield. We also evaluated the use of ancestry-stratified allele frequency filters derived from an independent, diverse UK cohort (n = 33,724). FINDINGS: Compared with the European ancestry group, the East African group had nearly three times more variants prioritised for clinical review (IRR 2.77, 95% CI 2.33-3.29). Other non-European groups also had significantly higher counts. Diagnostic yield was similar across ancestry groups after adjustment (LRT p = 0.1650). Prioritised variants were less likely to be recorded as diagnostic in East African (OR 0.32, 95% CI 0.22-0.46), West African (0.47, 0.39-0.57), South Asian (0.65, 0.58-0.73), and Middle Eastern (0.68, 0.54-0.86) groups. Applying ancestry-stratified allele-frequency filters removed 3.1% of prioritised variants overall-24.3% in the East African group-without loss of diagnostic sensitivity, including 29.5% of recorded VUS in this group. INTERPRETATION: Differences in the likelihood of prioritised variants being recorded as diagnostic partly reflect limitations of current allele frequency resources, which use broad population groupings that mask within-group diversity. Increased representation of diverse ancestries in reference databases and better estimation of ancestry-appropriate allele frequencies will help reduce inefficiencies and improve equity in variant prioritisation for rare disease diagnosis. FUNDING: The UK Department of Health and Social Care and the EU's Horizon 2020 Research and Innovation Programme.

Humans↗

Partial sequence comparison of eight new Chinese strains of hepatitis E virus suggests the genome sequence is relatively stable.

Partial genomic sequences representing 420 nucleotides of a nonstructional region, 480 nucleotides of the putative RNA polymerase region, and 540 nucleotides of the structural region of epidemic-associated Chinese strains of hepatitis E virus (HEV) were obtained by direct sequencing of PCR-amplified DNA. Comparison with previously published HEV sequences showed a clear relatedness of all Chinese strains to each other and to a Pakistani strain (Sar-55). All eight Chinese strains examined had very similar sequences (98.5-99.8% homology) in the regions examined and were much closer to the Pakistani strain (Sar-55) (97.9-98.4% homology) than to the Burmese strain (92.5-93.3% homology). Sequence comparisons of the three genomic regions in the Chinese strains indicated that the RNA polymerase region was much more conserved than the other nonstructural region or the structural region. HEV isolates from three remote geographic regions of China had sequences closely related to each other.

Amino Acid Sequence↗

Partial genome sequencing of Rhodococcus equi ATCC 33701.

Preliminary analysis of a partial (30% coverage) genome sequence of Rhodococcus equi has revealed a number of important features. The most notable was the extent of the homology of genes identified with those of Mycobacterium tuberculosis. The similarities in the proportion of genes devoted to fatty acid degradation and to lipid biosynthesis was a striking but not surprising finding given the relatedness of these organisms and their success as intracellular pathogens. The rapid recent improvement in understanding of virulence in M. tuberculosis and other pathogenic mycobacteria has identified a large number of genes of putative or proven importance in virulence, homologs of many of which were also identified in R. equi. Although R. equi appears to have currently unique genes, and has important differences, its similarity to M. tuberculosis supports the need to understand the basis of virulence in this organism. The partial genome sequence will be a resource for workers interested in R. equi until such time as a full genome sequence has been characterized.

Aerobiosis↗