Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

The genome sequence of Sphagnum contortum Schultz, 1819 (Sphagnales: Sphagnaceae).

We present a genome assembly of Sphagnum contortum (twisted bog-moss; Streptophyta; Sphagnopsida; Sphagnales; Sphagnaceae). The genome sequence has a total length of 370.55 megabases. Most of the assembly (99.46%) is scaffolded into 21 chromosomal pseudomolecules. The mitochondrial sequence has a length of 141.7 kilobases and the plastid genome assembly has a length of 140.11 kilobases. This assembly was generated as part of the Darwin Tree of Life project, which produces reference genomes for eukaryotic species found in Britain and Ireland.

Sphagnales↗

The genome sequence of Gnaphalium uliginosum L., 1753 (Asterales: Asteraceae).

We present a genome assembly of Gnaphalium uliginosum (marsh cudweed; Streptophyta; Magnoliopsida; Asterales; Asteraceae). The genome sequence has a total length of 354.38 megabases. Most of the assembly (98.86%) is scaffolded into 7 chromosomal pseudomolecules. The mitochondrial sequence has a length of 195.0 kilobases and the plastid genome assembly has a length of 152.8 kilobases. Gene annotation of this assembly on Ensembl identified 24 149 protein-coding genes. This assembly was generated as part of the Darwin Tree of Life project, which produces reference genomes for eukaryotic species found in Britain and Ireland.

Asterales↗

RNA editing of rapeseed mitochondrial atp9 transcripts: RNA editing changes four amino acids, but termination codon is already encoded by genomic sequence.

The gene encoding subunit 9 of Fo-ATPase of rapeseed mitochondria has been isolated. The complete genomic DNA sequence and cDNA sequence corresponding to the atp9 gene transcript have been determined by a method involving cDNA synthesis, using specific oligonucleotides as primers, followed by PCR amplification, cloning and sequencing of the amplification products. In comparison of cDNA sequences to genomic one, four modifications, C-to-U conversions, have been found. When compared with RNA editing patterns of atp9 transcripts among plant mitochondria, that of rapeseed atp9 transcript is more simple; there are only four editing sites on the coding region, and its termination codon is already encoded by genomic sequence.

Amino Acid Sequence↗

The genome sequence of Cardamine flexuosa With., 1796 (Brassicales: Brassicaceae).

We present a genome assembly of Cardamine flexuosa (Wavy Bitter-cress; Streptophyta; Magnoliopsida; Brassicales; Brassicaceae). The genome sequence has a total length of 204.54 megabases. Most of the assembly (97.53%) is scaffolded into 8 chromosomal pseudomolecules. The mitochondrial sequence has a length of 299.98 kilobases and the plastid genome assembly has a length of 153.92 kilobases. Gene annotation of this assembly on Ensembl identified 24 305 protein-coding genes. This assembly was generated as part of the Darwin Tree of Life project, which produces reference genomes for eukaryotic species found in Britain and Ireland.

Brassicales↗

The genome sequence of Veronica verna L., 1753 (Lamiales: Plantaginaceae).

We present a genome assembly of Veronica verna (Spring Speedwell; Streptophyta; Magnoliopsida; Lamiales; Plantaginaceae). The genome sequence has a total length of 463.88 megabases. Most of the assembly (98.85%) is scaffolded into 8 chromosomal pseudomolecules. The mitochondrial sequence has a length of 311.81 kilobases and the plastid genome assembly has a length of 149.82 kilobases. Gene annotation of this assembly on Ensembl identified 22 903 protein-coding genes. This assembly was generated as part of the Darwin Tree of Life project, which produces reference genomes for eukaryotic species found in Britain and Ireland.

Asterales↗

Draft genome sequences of four strains of Ralstonia pseudosolanacearum.

We report the draft genome sequences of four Ralstonia pseudosolanacearum strains obtained from the NARO Genebank, Japan. Illumina MiSeq i100 sequencing generated 2 × 300 bp paired-end reads. The draft assemblies ranged from 5.67 to 5.86 Mb, with GC contents of 66.9%-67.0%.

Ralstonia pseudosolanacearum↗

Findings emerging from complete microbial genome sequences.

Sixteen microorganisms, including one eukaryote, four archaeons, and 11 eubacteria, have been completely sequenced and published. More than 50 genomes are scheduled to be completed by the year 2000. This explosive growth of information is forcing change in many scientific disciplines (e.g. bioinformatics and molecular genetics), spawning new fields, and even changing the way scientific information is used and shared. Novel, global genome sequence comparisons seem slow to appear but the infrastructure for these projects is being built, and we expect exciting developments in the near future.

Databases, Factual↗

Phylogenetic analyses of Vitis (Vitaceae) based on complete chloroplast genome sequences: effects of taxon sampling and phylogenetic methods on resolving relationships among rosids.

BACKGROUND: The Vitaceae (grape) is an economically important family of angiosperms whose phylogenetic placement is currently unresolved. Recent phylogenetic analyses based on one to several genes have suggested several alternative placements of this family, including sister to Caryophyllales, asterids, Saxifragales, Dilleniaceae or to rest of rosids, though support for these different results has been weak. There has been a recent interest in using complete chloroplast genome sequences for resolving phylogenetic relationships among angiosperms. These studies have clarified relationships among several major lineages but they have also emphasized the importance of taxon sampling and the effects of different phylogenetic methods for obtaining accurate phylogenies. We sequenced the complete chloroplast genome of Vitis vinifera and used these data to assess relationships among 27 angiosperms, including nine taxa of rosids. RESULTS: The Vitis vinifera chloroplast genome is 160,928 bp in length, including a pair of inverted repeats of 26,358 bp that are separated by small and large single copy regions of 19,065 bp and 89,147 bp, respectively. The gene content and order of Vitis is identical to many other unrearranged angiosperm chloroplast genomes, including tobacco. Phylogenetic analyses using maximum parsimony and maximum likelihood were performed on DNA sequences of 61 protein-coding genes for two datasets with 28 or 29 taxa, including eight or nine taxa from four of the seven currently recognized major clades of rosids. Parsimony and likelihood phylogenies of both data sets provide strong support for the placement of Vitaceae as sister to the remaining rosids. However, the position of the Myrtales and support for the monophyly of the eurosid I clade differs between the two data sets and the two methods of analysis. In parsimony analyses, the inclusion of Gossypium is necessary to obtain trees that support the monophyly of the eurosid I clade. However, maximum likelihood analyses place Cucumis as sister to the Myrtales and therefore do not support the monophyly of the eurosid I clade. CONCLUSION: Phylogenies based on DNA sequences from complete chloroplast genome sequences provide strong support for the position of the Vitaceae as the earliest diverging lineage of rosids. Our phylogenetic analyses support recent assertions that inadequate taxon sampling and incorrect model specification for concatenated multi-gene data sets can mislead phylogenetic inferences when using whole chloroplast genomes for phylogeny reconstruction.

Base Sequence↗

TMBETA-GENOME: database for annotated beta-barrel membrane proteins in genomic sequences.

We have developed the database, TMBETA-GENOME, for annotated beta-barrel membrane proteins in genomic sequences using statistical methods and machine learning algorithms. The statistical methods are based on amino acid composition, reside pair preference and motifs. In machine learning techniques, the combination of amino acid and dipeptide compositions has been used as main attributes. In addition, annotations have been made using the criterion based on the identification of beta-barrel membrane proteins and exclusion of globular and transmembrane helical proteins. A web interface has been developed for identifying the annotated beta-barrel membrane proteins in all known genomes. The users have the feasibility of selecting the genome from the three kingdoms of life, archaea, bacteria and eukaryote, and five different methods. Further, the statistics for all genomes have been provided along with the links to different algorithms and related databases. It is freely available at http://tmbeta-genome.cbrc.jp/annotation/.

Algorithms↗

A picorna-like virus from the red imported fire ant, Solenopsis invicta: initial discovery, genome sequence, and characterization.

We report the first discovery and genome sequence of a virus infecting the red imported fire ant, Solenopsis invicta. The 8026 nucleotide, polyadenylated, RNA genome encoded two large open reading frames (ORF1 and ORF2), flanked and separated by 27, 223, and 171 nucleotide untranslated regions, respectively. The predicted amino acid sequence of the 5' proximal ORF1 (nucleotides 28 to 4218) exhibited significant identity and possessed consensus sequences characteristic of the helicase, cysteine protease, and RNA-dependent RNA polymerase sequence motifs from picornaviruses, picorna-like viruses, comoviruses, caliciviruses, and sequiviruses. The predicted amino acid sequence of the 3' proximal ORF2 (nucleotides 4390-7803) showed similarity to structural proteins in picorna-like viruses, especially the acute bee paralysis virus. Electron microscopic examination of negatively stained samples from virus-infected fire ants revealed isometric particles with a diameter of 31 nm, consistent with Picornaviridae. A survey for the fire ant virus from areas around Florida revealed a pattern of fairly widespread distribution. Among 168 nests surveyed, 22.9% were infected. The virus was found to infect all fire ant caste members and developmental stages, including eggs, early (1st-2nd) and late (3rd-4th) instars, worker pupae, workers, sexual pupae, alates ( male symbol and female symbol ), and queens. The virus, tentatively named S. invicta virus (SINV-1), appears to belong to the picorna-like viruses. We did not observe any perceptible symptoms among infected nests in the field. However, in every case where an SINV-1-infected colony was excavated from the field with an inseminated queen and held in the laboratory, all of the brood in these colonies died within 3 months.

3' Untranslated Regions↗

Completion of the sequence of a cetacean morbillivirus and comparative analysis of the complete genome sequences of four morbilliviruses.

The gene encoding the large (L) protein and the genome termini of the dolphin strain of cetacean morbillivirus (CeMV) were sequenced. The CeMV genome is 15702 nucleotides long and has been compared with other available morbillivirus genome sequences in regards to the "rule of six" and the "phase" of any particular nucleotide, defined as its position within a given hexamer, which here is defined as a group of six nucleotides starting from the 3' end of the genomic RNA. With exception of the position of the start of the F gene, the phase of the transcription start sites of each gene is strictly conserved between the morbilliviruses, but each gene is in a different phase. The lengths of gene transcripts differ between viruses by multiples of six nucleotides with exception of the M and F transcripts. The differences between the various morbilliviruses result from deletions or insertions of multiples of six nucleotides in the 3' and 5' UTRs of the different viral genes. The four bases were distributed non-randomly over the six positions in the hexamer boxes. However, the distribution patterns of each of the four bases indicated that multiples of three were more prevalent than those of six nucleotides. This reflected the positions of nucleotides in codons and codon usage in the reading frames. The L protein of CeMV was found to be 2183 amino acids in length and similar to that of MV and RPV. The CeMV L protein sequence was found to be equidistant between those of the CDV/PDV and MV/RPV subgroups of the morbilliviruses. This concurs with the analyses carried out on the other structural proteins.

3' Untranslated Regions↗

Complete genomic sequences for hepatitis C virus subtypes 6e and 6g isolated from Chinese patients with injection drug use and HIV-1 co-infection.

In one of our recent studies, two HCV genotype 6 variants were identified in patients from Hong Kong and Guangxi in southern China, with injection drug use and HIV-1 co-infection. We report the complete genomic sequences for these two variants: GX004 and HK6554. Their entire genome lengths were 9,468 and 9,462 nt; the 5' UTRs were 338 nt followed by single ORFs of 9,069 nt; the 3' UTRs were 61 and 55 nt including 29 and 23 nt poly(U) tracks. Phylogenetic analysis using a maximum likelihood method showed that HK6554 was classified into subtype 6g and GX004 represented the first complete genome sequence for subtype 6e. Further analysis with reference sequences in three different genomic regions revealed that GX004 closely clustered with a group of subtype 6e variants, which were previously exclusively found in Vietnam and recently increasingly identified in injection drug users from the Guangxi province in southern China that borders Vietnam. This suggests that subtype 6e could become epidemic in southern China by network transmission among injection drug users.

China↗

Exploiting genome sequence: predictions for mechanisms of Campylobacter chemotaxis.

The genome sequence of Campylobacter jejuni NCTC 11168 reveals the presence of orthologues of the chemotaxis genes cheA, cheW, cheV, cheY, cheR and cheB, ten chemoreceptor genes and two aerotaxis genes. The presence of cheV and a response regulator domain in CheA, combined with the absence of a cheZ gene and the lack of a response regulator domain in CheB, reveals significant differences in the C. jejuni chemotaxis system compared with that found in other bacteria.

Bacterial Proteins↗

The genome sequence of Prunus padus L., 1753 (Rosales: Rosaceae).

We present a genome assembly of Prunus padus (bird cherry; Streptophyta; Magnoliopsida; Rosales; Rosaceae). The genome sequence has a total length of 454.49 megabases. Most of the assembly (97.36%) is scaffolded into 16 chromosomal pseudomolecules. The mitochondrial sequences have lengths of 280.28 and 155.92 kilobases and the plastid genome assembly has a length of 158.96 kilobases. This assembly was generated as part of the Darwin Tree of Life project, which produces reference genomes for eukaryotic species found in Britain and Ireland.

Prunus padus↗

The genome sequence of Quercus cerris L., 1753 (Fagales: Fagaceae).

We present a genome assembly of Quercus cerris (Turkey oak; Streptophyta; Magnoliopsida; Fagales; Fagaceae). The genome sequence has a total length of 777.39 megabases. Most of the assembly (99.89%) is scaffolded into 12 chromosomal pseudomolecules. The mitochondrial sequences have lengths of 380.47 and 15.0 kilobases and the plastid genome assembly has a length of 161.2 kilobases. Gene annotation of this assembly on Ensembl identified 25 805 protein-coding genes. This assembly was generated as part of the Darwin Tree of Life project, which produces reference genomes for eukaryotic species found in Britain and Ireland.

Fagales↗

Human interleukin-9: genomic sequence, chromosomal location, and sequences essential for its expression in human T-cell leukemia virus (HTLV)-I-transformed human T cells.

We have isolated the genomic sequence of human interleukin-9 (IL-9) based on its sequence homology with a human IL-9 cDNA isolated from human T-cell leukemia virus (HTLV)-I-transformed T cells by expression cloning. The entire genomic sequence has been determined and the gene consists of five exons and four introns. The human IL-9 gene is mapped to the long arm of human chromosome 5 at band 5q31-32, a region found to be deleted in a number of patients with acquired 5q- abnormalities and hematologic disorders. Several blocks of transcriptional control sequences have been identified at the 5'-flanking region of the human IL-9 gene that may play an important role in the control of IL-9 gene expression. The 5'-regulatory region of the human IL-9 gene also contains sequences identified in the 5'-flanking regions of other cytokine genes mapped to the long arm of human chromosome 5, including IL-3, IL-4, IL-5, and granulocyte-macrophage colony-stimulating factor and other T-cell growth factor genes including IL-2 and IL-6. The IL-9 gene is constitutively expressed in the HTLV-I-transformed human T cells and the expression of IL-9 in these cells can be further induced by 12-O-tetradecanoyl phorbol 13-acetate. Transient transfection analysis using the plasmid containing the 5'-flanking region of IL-9 gene upstream from the firefly luciferase ciferase report gene indicated that the 0.9-kb Smal-Sacl fragment of the IL-9 gene contains sequences required for the constitutive and activated expression of IL-9 gene in HTLV-I-transformed cells. These results will now allow us to study the regulatory mechanism of IL-9 gene expression in normal and leukemic human T cells.

Amino Acid Sequence↗

AF4/FEL, a gene involved in infant leukemia: sequence variations, gene structure, and possible homology with a genomic sequence on 5q31.

The most common chromosome abnormality among infants with acute lymphoblastic leukemia is a t(4;11)(q2l;q23) and patients with this 4;11 translocation have a very poor prognosis. This unique genetic rearrangement fuses the MLL/ALL-1/HRX-Htrx gene at 11q23 with the AF4/FEL gene at 4q21. The resulting chimeric mRNAs presumably encode chimeric proteins which contribute to the leukemogenic state. The AF4 gene remains poorly understood with an unknown function. In this report, we describe the cDNA sequence information from human placental tissue where AF4 mRNA is highly expressed. We identified six intron-exon boundaries in the AF4 genomic structure and discussed more than 30 AF4 cDNA sequence variations reported in the literature. In addition, we identified three overlapping genomic sequences in GenBank entitled the "interleukin growth hormone cluster on chromosome 5q31," which, when aligned and translated, had three regions that suggested homology to the predicted AF4 protein sequence (32% amino acid sequence identity over 314 amino acids, 43% over 63 amino acids, and 50% over 40 amino acids). Of interest, this same chromosome 5q31 region has also been implicated in MLL gene rearrangements in human leukemia.

Amino Acid Sequence↗

Whole genome sequence of a superbug-Escherichia coli strain KAB-AI-497 isolated from the vagina of a 20 year old pregnant woman with premature rupture of membrane (PROM) in a resource limited setting, Kabale Regional Referral Hospital, in Uganda.

OBJECTIVES: The objective of the study is to sequence the whole genome of multidrug resistant E. coli strain KAB-AI-497 that causes bacterial vaginosis and implicated in premature rupture of membrane in pregnant woman. DATA DESCRIPTION: The DNA of the E. coli strain KAB-AI-497 was extracted using the MagAttract HMW DNA Kit, and the extracted DNA was sequenced using an MGI DNBSEQ G99ARS platform. FastQC was used to perform quality control analysis and the reads were trimmed by Trimmomatic. De novo genome assembly was performed by SPAdes and it resulted to a draft assembled genome that has 5.1 Mb genome size, 153 contigs, and 50.5% GC content. Quality analysis of the assembled genome revealed it has 98.46% completeness and 0.97% contamination. The closest E. coli strain to this strain KAB-AI-497 in terms of similarity was Escherichia coli SMS-3-5 with an average nucleotide identity of 98.43% and genome coverage of 86.19%, which confirmed the species level identity of the strain. The assembled genome was annotated using the NCBI Prokaryotic Genome Annotation Pipeline which identified 4,726 protein coding genes in the strain genome. Furthermore, the annotation revealed the genome has resistant genes responsible for resistance against many antibiotic classes such as tetracycline, fluoroquinolone, and penicillin.

Escherichia coli↗