Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Isolation of the genome sequence strain Mycobacterium avium 104 from multiple patients over a 17-year period.

The genome sequence strain 104 of the opportunistic pathogen Mycobacterium avium was isolated from an adult AIDS patient in Southern California in 1983. Isolates of non-paratuberculosis M. avium from 207 other patients in Southern California and elsewhere were examined for genotypic identity to strain 104. This process was facilitated by the use of a novel two-step approach. In the first step, all 208 strains in the sample were subjected to a high-throughput, large sequence polymorphism (LSP)-based genotyping test, in which DNA from each strain was tested by PCR for the presence or absence of 4 hypervariable genomic regions. Nineteen isolates exhibited an LSP type that resembled that of strain 104. This subset of 19 isolates was then subjected to high-resolution repetitive sequence-based PCR typing, which identified 10 isolates within the subset that were genotypically identical to strain 104. These isolates came from 10 different patients at 5 clinical sites in the western United States, and they were isolated over a 17-year time span. Therefore, the sequenced genome of M. avium strain 104 has been associated with disease in multiple patients in the western United States. Although M. avium is known for its genetic plasticity, these observations also show that strains of the pathogen can be genotypically stable over extended time periods.

Bacterial Typing Techniques↗

Genomic sequence analysis tools: a user's guide.

The wealth of information from various genome sequencing projects provides the biologist with a new perspective from which to analyze, and design experiments with, mammalian systems. The complexity of the information, however, requires new software tools, and numerous such tools are now available. Which type and which specific system is most effective depends, in part, upon how much sequence is to be analyzed and with what level of experimental support. Here we survey a number of mammalian genomic sequence analysis systems with respect to the data they provide and the ease of their use. The hope is to aid the experimental biologist in choosing the most appropriate tool for their analyses.

Internet↗

Full genome sequence and analysis of Indian swine hepatitis E virus isolate of genotype 4.

The full-length genomic sequence of an Indian swine hepatitis E virus (HEV) isolate (IND-SW-00-01) recovered from feces of a pig experimentally infected with swine HEV pool from western India was determined. The genome consisted of 7,240 nucleotides, excluding the poly (A) tail of at least 22 residues and contained three open reading frames (ORFs), ORF-1 encoding 1,707 amino acids, ORF-2 encoding 674 amino acids and ORF-3 encoding 114 amino acids. Comparative full-length genome sequence and phylogenetic analyses suggested that the Indian swine HEV represents a distinct variant among the genotype 4 isolates with a divergence of 15-16.6%. Analyses based on ORF-1, 2 and 3 as well as partial ORF-2 (227 nucleotides) yielded similar results. As compared to type 4 HEV isolates, 26 unique amino acid substitutions were recorded, 16 in ORF-1, 8 in ORF-2 and 2 in ORF-3. IND-SW-00-01 showed insertion of 'C' at 5159 position while all other type 4 isolates have insertion of 'U' at the same position. Whether these changes contribute towards observed absence of type 4 HEV infections in Indian patients needs to be determined.

3' Untranslated Regions↗

An interactive tool for extracting exons and SNP from genomic sequence: isolation of HCN1 and HCN3 ion channel genes.

Genome Analyzer (GenoA) with a relational database back-end, was developed to extract information from mammalian genomic sequences. This data mining and visualization tool-set enables laboratory bench scientists to identify and assemble virtual cDNA from genomic exon sequences, and provides a starting point to identify potential alternative splice variants and polymorphisms in silico. The study described in this paper demonstrates the use of GenoA to study human brain hyperpolarization-activated cation channel genes HCN1 and HCN3.

Amino Acid Sequence↗

Short report: genetic heterogeneity of Japanese encephalitis virus assessed via analysis of the full-length genome sequence of a Korean isolate.

We determined the full-length genome sequence of Japanese encephalitis virus (JEV) K94P05 isolated in Korea. Sequence analysis showed that the 10,963-nucleotide-long RNA genome of K94P05 was 13 or 14 nucleotides shorter than the genome of other JEV isolates because of a deletion in the 3' noncoding region of K94P05. Compared with sequences of other JEV isolates, the full-length nucleotide sequence showed 89.0-89.6% homology, and the deduced amino acid sequence showed between 96.4-97.3% homology. A region of approximately 60 nucleotides immediately downstream of the open reading frame stop codon of K94P05 showed high sequence variability as compared with other JEV isolates. K94P05 formed a distinct group within a phylogenetic tree established with the full-length genome sequences. Cross-neutralization studies showed that polyclonal antibodies to Korean isolates were 3 times better at neutralizing the Korean isolates than antibodies to Nakayama-NIH. These findings suggest that Korean JEV K94P05 is genetically and antigenically distinct from other Asian JEV isolates.

Amino Acid Sequence↗

X-linked spondyloepiphyseal dysplasia tarda misdiagnosed as growth hormone deficiency: identification of a novel intronic TRAPPC2 variant by whole-genome sequencing.

BACKGROUND: X-linked spondyloepiphyseal dysplasia tarda (SEDT) is a rare skeletal dysplasia caused by pathogenic variants in TRAPPC2 and typically presents in late childhood or adolescence with short-trunk disproportion and vertebral dysplasia. CASE PRESENTATION: We describe a family series centered on an adolescent male initially diagnosed with GHD due to reduced height velocity and subnormal GH stimulation results, who received recombinant human GH (rhGH) therapy for three years with negligible improvement. During puberty, he developed progressive short-trunk disproportion and characteristic radiographic features, including platyspondyly and posterior hump-shaped vertebral endplates, suggestive of SEDT. Whole-exome sequencing (WES) was nondiagnostic, whereas whole-genome sequencing (WGS) identified a novel intronic TRAPPC2 variant, c.239-20_239-12delinsAATGAA, initially classified as a variant of uncertain significance (VUS). Segregation analysis across the family enabled reclassification of the variant to likely pathogenic, confirming X-linked SEDT. The proband's younger brother exhibited earlier radiologic abnormalities and, notably, a favorable response to rhGH, whereas the younger sister-an asymptomatic heterozygous carrier-showed normal spinal morphology, consistent with expected female carrier phenotypes. CONCLUSIONS: This family-based report underscores the generally limited therapeutic effect of rhGH in SEDT while highlighting potential interindividual variability, as evidenced by the younger male sibling's response. It further emphasizes the diagnostic utility of WGS for detecting deep intronic variants missed by WES and the importance of segregation analysis in resolving VUS in rare skeletal dysplasias.

Humans↗

A deep intronic IFT172 variant causing pseudoexon inclusion identified by whole-genome sequencing in nephronophthisis.

Nephronophthisis is an autosomal recessive ciliopathy and a major genetic cause of end-stage kidney disease in children and young adults. Although next-generation sequencing panels have improved diagnostic yield, some patients remain genetically unresolved, partly due to deep intronic variants that disrupt pre-mRNA splicing and are not captured by exon-focused approaches. We report a 13-year-old boy who presented with advanced kidney dysfunction, small renal cysts, and kidney histopathology consistent with nephronophthisis. Targeted gene panel sequencing failed to identify causative pathogenic variants beyond a missense variant of uncertain significance. Whole-genome sequencing subsequently revealed compound heterozygous variants in IFT172 (NM_015662.3): a missense variant (c.4696C > T, p.Arg1566Cys) and a deep intronic variant (c.4915-94A > G). In silico analysis predicted activation of cryptic splice sites leading to inclusion of an 86-bp pseudoexon, which was confirmed by a minigene splicing assay. These findings established a molecular diagnosis of IFT172-related nephronophthisis. To our knowledge, this is the first report demonstrating pseudoexon inclusion in IFT172, thereby expanding its mutational spectrum. Our case underscores the importance of evaluating deep intronic regions using whole-genome sequencing and functional validation in genetically unresolved nephronophthisis.

Humans↗

Chicken leukosis virus genome sequences in DNA from normal chick cells and virus-induced bursal lymphomas.

Genome sequences of two recent field isolates of avian leukosis viruses in the DNA of normal and neoplastic chicken cells were studied by DNA-RNA hybridization under conditions of DNA excess. Comparisons were made between 60-70S RNA from these viruses and that of a chicken endogenous type C virus (RAV-0), and of a series of "laboratory" leukosis and sarcoma viruses, by competitive hybridization analysis. A minimum of 18% of the genome sequences of both ALV isolates detected in DNA from lymphomas they induced were not detected in normal chicken DNA. The vast majority of the fraction of RNA sequences from ALV which do form hybrids with normal chick DNA appear to be reacting with the endogenous provirus of RAV-0. The genomic representation of a variety of avian leukosis and sarcoma viruses in normal chicken cells could not be distinguished by these methods (except that 13% of the RAV-0 genome was not shared with any of the other viruses). In contrast, the portion of the ALV genome exogenous to the normal chicken geome showed significant divergence from that of two sarcoma viruses (Pr RSV-C and B-77). The increased hybridization of ALV RNA with lymphoma DNA was used to detect the appearance of ALV specific sequences in the bursa of Fabricius following infection.increased hybridization was correlated with both the time after infection and the extent of replacement of the bursa by lymphoma. About one half of the increase in hybridization preceded histologic evidence of transformation.

Alpharetrovirus↗

Prediction of probable genes by Fourier analysis of genomic sequences.

MOTIVATION: The major signal in coding regions of genomic sequences is a three-base periodicity. Our aim is to use Fourier techniques to analyse this periodicity, and thereby to develop a tool to recognize coding regions in genomic DNA. RESULT: The three-base periodicity in the nucleotide arrangement is evidenced as a sharp peak at frequency f = 1/3 in the Fourier (or power) spectrum. From extensive spectral analysis of DNA sequences of total length over 5.5 million base pairs from a wide variety or organisms (including the human genome), and by separately examining coding and non-coding sequences, we find that the relative-height of the peak at f = 1/3 in the Fourier spectrum is a good discriminator of coding potential. This feature is utilized by us to detect probable coding regions in DNA sequences, by examining the local signal-to-noise ratio of the peak within a sliding window. While the overall accuracy is comparable to that of other techniques currently in use, the measure that is presently proposed is independent of training sets or existing database information, and can thus find general application. AVAILABILITY: A computer program GeneScan which locates coding open reading frames and exonic regions in genomic sequences has been developed, and is available on request.

Algorithms↗

Mutation analysis of 20 SARS virus genome sequences: evidence for negative selection in replicase ORF1b and spike gene.

AIM: Recently, more SARS-CoV virus genome sequences are released to the GenBank database. The aim of this study is to reveal the evolution forces of SARS-CoV virus by analyzing the nucleotide mutations in these sequences. METHODS: We obtained 20 SARS-CoV virus genome sequences from NCBI database, and calculated the ratio of non-synonymous nucleotide substitution per non-synonymous site (Ka) and synonymous nucleotide substitution per synonymous site (Ks) for SARS-CoV virus genes. RESULTS: The Ka/Ks ratios for replicase polyprotein ORF1a, ORF1b, and spike protein gene are 1.09 (P=0.6501), 0.38 (P=0.0074), 0.65 (P=0.0685) respectively. CONCLUSION: SARS-CoV virus replicase polyprotein ORF1b is undergoing negative selection; negative selection force is also probably operating on spike protein gene. These results provide basis for future developing a new drug and vaccine against SARS.

Base Sequence↗

Nucleotide sequence of the 5'-terminus of Newcastle disease virus and assembly of the complete genomic sequence: agreement with the "rule of six".

We have determined the sequences of the 5' ends of three strains of Newcastle disease virus, permitting the assembly of the entire genomic sequence, which amounts to 15,186 nucleotides. This length is in agreement with the rule of six, which has been shown to determine replication efficiency in similar viruses. Comparison of the extreme 5' end of the trailer sequence with that of the 3'-terminal leader sequence of the virus reveals a high degree of complementarity. Variation between the 5'-terminal sequences of the different strains reveals the presence of alternative L gene polyadenylation signals, leading to correspondingly different trailer lengths.

Animals↗

Complete genome sequence of Methanobacterium thermoautotrophicum deltaH: functional analysis and comparative genomics.

The complete 1,751,377-bp sequence of the genome of the thermophilic archaeon Methanobacterium thermoautotrophicum deltaH has been determined by a whole-genome shotgun sequencing approach. A total of 1,855 open reading frames (ORFs) have been identified that appear to encode polypeptides, 844 (46%) of which have been assigned putative functions based on their similarities to database sequences with assigned functions. A total of 514 (28%) of the ORF-encoded polypeptides are related to sequences with unknown functions, and 496 (27%) have little or no homology to sequences in public databases. Comparisons with Eucarya-, Bacteria-, and Archaea-specific databases reveal that 1,013 of the putative gene products (54%) are most similar to polypeptide sequences described previously for other organisms in the domain Archaea. Comparisons with the Methanococcus jannaschii genome data underline the extensive divergence that has occurred between these two methanogens; only 352 (19%) of M. thermoautotrophicum ORFs encode sequences that are >50% identical to M. jannaschii polypeptides, and there is little conservation in the relative locations of orthologous genes. When the M. thermoautotrophicum ORFs are compared to sequences from only the eucaryal and bacterial domains, 786 (42%) are more similar to bacterial sequences and 241 (13%) are more similar to eucaryal sequences. The bacterial domain-like gene products include the majority of those predicted to be involved in cofactor and small molecule biosyntheses, intermediary metabolism, transport, nitrogen fixation, regulatory functions, and interactions with the environment. Most proteins predicted to be involved in DNA metabolism, transcription, and translation are more similar to eucaryal sequences. Gene structure and organization have features that are typical of the Bacteria, including genes that encode polypeptides closely related to eucaryal proteins. There are 24 polypeptides that could form two-component sensor kinase-response regulator systems and homologs of the bacterial Hsp70-response proteins DnaK and DnaJ, which are notably absent in M. jannaschii. DNA replication initiation and chromosome packaging in M. thermoautotrophicum are predicted to have eucaryal features, based on the presence of two Cdc6 homologs and three histones; however, the presence of an ftsZ gene indicates a bacterial type of cell division initiation. The DNA polymerases include an X-family repair type and an unusual archaeal B type formed by two separate polypeptides. The DNA-dependent RNA polymerase (RNAP) subunits A', A", B', B" and H are encoded in a typical archaeal RNAP operon, although a second A' subunit-encoding gene is present at a remote location. There are two rRNA operons, and 39 tRNA genes are dispersed around the genome, although most of these occur in clusters. Three of the tRNA genes have introns, including the tRNAPro (GGG) gene, which contains a second intron at an unprecedented location. There is no selenocysteinyl-tRNA gene nor evidence for classically organized IS elements, prophages, or plasmids. The genome contains one intein and two extended repeats (3.6 and 8.6 kb) that are members of a family with 18 representatives in the M. jannaschii genome.

Anaerobiosis↗

The genomic sequence of cardamine chlorotic fleck carmovirus.

The complete genomic sequence of cardamine chlorotic fleck carmovirus (CCFV) has been determined. The genome is a positive-sense ssRNA molecule 4041 nucleotides in length, and has 47 to 64% sequence identity with turnip crinkle, carnation mottle and melon necrotic spot carmoviruses. CCFV and these other carmoviruses have four similar open reading frames (ORFs), and CCFV has large regions of amino acid identity in all of these ORFs with a European isolate of turnip crinkle virus. CCFV, which replicates well in Arabidopsis thaliana, has only been found so far in Australia in the wild perennial brassica Cardamine lilacina.

Amino Acid Sequence↗

Complete genome sequence of the grouper iridovirus and comparison of genomic organization with those of other iridoviruses.

The complete DNA sequence of grouper iridovirus (GIV) was determined using a whole-genome shotgun approach on virion DNA. The circular form genome was 139,793 bp in length with a 49% G + C content. It contained 120 predicted open reading frames (ORFs) with coding capacities ranging from 62 to 1,268 amino acids. A total of 21% (25 of 120) of GIV ORFs are conserved in the other five sequenced iridovirus genomes, including DNA replication, transcription, nucleotide metabolism, protein modification, viral structure, and virus-host interaction genes. The whole-genome nucleotide pairwise comparison showed that GIV virus was partially colinear with counterparts of previously sequenced ranaviruses (ATV and TFV). Besides, sequence analysis revealed that GIV possesses several unique features which are different from those of other complete sequenced iridovirus genomes: (i) GIV is the first ranavirus-like virus which has been sequenced completely and which infects fish other than amphibians, (ii) GIV is the only vertebrate iridovirus without CpG sequence methylation and lacking DNA methyltransferase, (iii) GIV contains a purine nucleoside phosphorylase gene which is not found in other iridoviruses or in any other viruses, (iv) GIV contains 17 sets of repeat sequence, with basic unit sizes ranging from 9 to 63 bp, dispersed throughout the whole genome. These distinctive features of GIV further extend our understanding of molecular events taking place between ranavirus and its hosts and the iridovirus evolution.

Animals↗

[Genomic sequence of hepatitis A virus L-A-1 vaccine strain].

OBJECTIVE: To study the genome sequence of hepatitis A virus L-A-1 strain which has been applied for live attenuated vaccine production in China, to compare with other HAV strains, to understand some characteristics of L-A-1 strain, and to find the mechanism of attenuation and cell adaptation. METHODS: Genome fragments were prepared by antigen-capture PCR from infected cell (2BS), PCR products were cloned into T vector, sequenced and analyzed by using bioinformatics program. RESULTS: Analysis of the genomic sequences(nt 25-7,418) showed that the open reading frame contains 6,675 nucleotides in length encoding 2,225 amino acids. Sequence homology comparison showed 98.00% and 94.00% homology at nucleotide level, and 98.51% and 98.65% homology at amino acid level with international strains MBB and HM 175, respectively. Through comparison with other attenuated, cell adapted and cytopathic effect (CPE) strains, L-A-1 strain had mutation at nt 152, 591, 646, 687 and insertion at nt 180-181 in 5?NTR and had mutation at nt 3,889 (aa 1 052-Val) in 2B region, these mutations and insertion are molecular basis for cell adaptation; mutation at nt 4,185 (aa 1 152-Lys) in 2C region should be attenuated marker; deletion in 3A region (nt 5,020-5,025) that caused two amino acids deletion is virus fast growth basis. CONCLUSION: Through analyzing L-A-1 strain genomic sequence, certain sites related to cell adaptation and attenuation were found.

Adaptation, Biological↗

Gene discovery through genomic sequencing of Brucella abortus.

Brucella abortus is the etiological agent of brucellosis, a disease that affects bovines and human. We generated DNA random sequences from the genome of B. abortus strain 2308 in order to characterize molecular targets that might be useful for developing immunological or chemotherapeutic strategies against this pathogen. The partial sequencing of 1,899 clones allowed the identification of 1,199 genomic sequence surveys (GSSs) with high homology (BLAST expect value < 10(-5)) to sequences deposited in the GenBank databases. Among them, 925 represent putative novel genes for the Brucella genus. Out of 925 nonredundant GSSs, 470 were classified in 15 categories based on cellular function. Seven hundred GSSs showed no significant database matches and remain available for further studies in order to identify their function. A high number of GSSs with homology to Agrobacterium tumefaciens and Rhizobium meliloti proteins were observed, thus confirming their close phylogenetic relationship. Among them, several GSSs showed high similarity with genes related to nodule nitrogen fixation, synthesis of nod factors, nodulation protein symbiotic plasmid, and nodule bacteroid differentiation. We have also identified several B. abortus homologs of virulence and pathogenesis genes from other pathogens, including a homolog to both the Shda gene from Salmonella enterica serovar Typhimurium and the AidA-1 gene from Escherichia coli. Other GSSs displayed significant homologies to genes encoding components of the type III and type IV secretion machineries, suggesting that Brucella might also have an active type III secretion machinery.

Brucella abortus↗

New perspectives on rickettsial evolution from new genome sequences of rickettsia, particularly R. canadensis, and Orientia tsutsugamushi.

The complete genome sequences available for eight species of Rickettsia and information for other near relatives in the Rickettsiales including Orientia and species of Anaplasmataceae are a rich resource for comparative analyses of the evolution of these obligate intracellular bacteria. Differences in these organisms have permitted them to colonize varied intracellular compartments, arthropod vectors, and vertebrate reservoirs in both pathogenic and symbiotic relationships. We summarize some comparative aspects of the genomes of these organisms, paying particular attention to the recently completed sequence for R. canadensis McKiel strain and an estimated two-thirds of the genome sequence for a Thailand patient isolate of Orientia tsutsugamushi. The Rickettsia genomes exhibit a high degree of synteny punctuated by distinctive chromosome inversions and consistent phylogenetic relationships regardless of whether protein coding sequences or RNA genes, concatenated open reading frames or gene regions, or whole genomes are used to construct phylogenetic trees. The aggregate characteristics (number, length, composition, repeat identity) of tandem repeat sequences of Rickettsia, which often exhibit recent and rapid divergence between closely related strains and species of bacteria, are also very conserved in Rickettsia but differed significantly in Orientia. O. tsutsugamushi shared no significant synteny to species of Rickettsia or Anaplasmataceae, supporting its placement in a unique genus. Like Rickettsia felis, Orientia has many transposases and ankyrin and tetratricopeptide repeat domains. Orientia shares the important ATP/ADP translocase and proline-betaine transporter multigene families with Rickettsia, but has more gene families that may be involved in regulatory and transporter responses to environmental stimuli.

Animals↗

The complete genome sequence of Mycobacterium bovis.

Mycobacterium bovis is the causative agent of tuberculosis in a range of animal species and man, with worldwide annual losses to agriculture of $3 billion. The human burden of tuberculosis caused by the bovine tubercle bacillus is still largely unknown. M. bovis was also the progenitor for the M. bovis bacillus Calmette-Guérin vaccine strain, the most widely used human vaccine. Here we describe the 4,345,492-bp genome sequence of M. bovis AF2122/97 and its comparison with the genomes of Mycobacterium tuberculosis and Mycobacterium leprae. Strikingly, the genome sequence of M. bovis is >99.95% identical to that of M. tuberculosis, but deletion of genetic information has led to a reduced genome size. Comparison with M. leprae reveals a number of common gene losses, suggesting the removal of functional redundancy. Cell wall components and secreted proteins show the greatest variation, indicating their potential role in host-bacillus interactions or immune evasion. Furthermore, there are no genes unique to M. bovis, implying that differential gene expression may be the key to the host tropisms of human and bovine bacilli. The genome sequence therefore offers major insight on the evolution, host preference, and pathobiology of M. bovis.

Genome, Bacterial↗