Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 937 records · Page 52Linked to original sources

Hepatitis B virus genomes of patients with fulminant hepatitis do not share a specific mutation.

The pathogenesis of fulminant hepatitis B virus (HBV) infection is not well understood. The aim of this study was to investigate whether there is an association between specific viral variants and a fulminant disease course. The entire HBV genomes from the serum of eight patients with fulminant HBV infection and one patient with fulminant hepatitis during reinfection after liver transplantation were investigated. After isolation and amplification of viral DNA by polymerase chain reaction (PCR), plus and minus strands were directly sequenced. Sequence data were analyzed by comparative sequence alignments with 35 and 2 complete HBV genome sequences from patients without and with fulminant hepatitis, respectively. Several point mutations were present in all regions of the genomes. Many nucleotide changes had never or rarely been found in the reported HBV isolates from patients without fulminant hepatitis. A distinct mutation present in all genomes was not identified. Clusters of rare and unique mutations were observed in the enhancer II core promoter region. Mutations previously suggested to be associated with fulminant HBV infection were not consistently found. A precore stop codon mutation at nucleotide position 1896 or an A-to-T mutation at nucleotide position 1762 and a G-to-A mutation at nucleotide position 1764 in the core promoter region were present in four and three cases, respectively. Fulminant HBV infection does not appear to be caused by a specific genomic mutation. However, various mutations clustering in the enhancer II core promoter region may contribute to a fulminant disease course.

Adult↗

Phylogenetic relationships and evolutionary history of the reef fish family Labridae.

The family Labridae (including scarines and odacines) contains 82 genera and about 600 species of fishes that inhabit coastal and continental shelf waters in tropical and temperate oceans throughout the world. The Labridae (the wrasses) is the fifth largest fish family and second largest marine fish family, and is one of the most morphologically and ecologically diversified families of fishes in size, shape, and color. Labrid phylogeny is a long-standing problem in ichthyology that is part of the larger question of relationships within the suborder Labroidei. A phylogenetic analysis of labrids was conducted to investigate relationships among the six classical tribes of wrasses, the affinities of the wrasses to the parrotfishes (scarines), and the broad phylogenetic structure among labrid genera. Four gene fragments were sequenced from 98 fish species, including 84 labrid fishes and 14 outgroup taxa. Taxa were chosen from all major labrid clades and most major global ocean regions where labrid fishes exist, as well as cichlid, pomacentrid, and embiotocid outgroups. From the mitochondrial genome we sequenced portions of 12S rRNA (1000 bp) and 16S rRNA (585 bp), which were aligned by using a secondary structure model. From the nuclear genome, we sequenced part of the protein-coding genes RAG2 (846 bp) and Tmo4C4 (541 bp). Maximum likelihood, maximum parsimony, and Bayesian analyses on the resulting 2972 bp of DNA sequence produced similar topologies that confirm the monophyly of a family Labridae that includes the parrotfishes and butterfishes and strong support for many previously identified taxonomic subgroups. The tribe Hypsigenyini (hogfishes, tuskfishes) is the sister group to the remaining labrids and includes odacines and the chisel-tooth wrasse Pseudodax moluccanus, a species previously considered close to scarines. Cheilines and scarines are sister-groups, closely related to the temperate Labrini, and pseudocheilines and cheilines are split in all phylogenies. The razorfishes (novaculines) and temperate pseudolabrines form successive sister clades to the large crown group radiation of the Julidini. The cleaner wrasses (Labrichthyini) are nested within this radiation and several julidine genera do not appear to be monophyletic (e.g., Coris and Halichoeres). Invasion of temperate water by this predominantly tropical group has occurred multiple times and the reconstruction of biogeography assuming an Indo-Pacific ancestor results in five different lineages invading the Atlantic/Caribbean region. Functional novelties in the feeding apparatus have allowed labrid fishes to occupy nearly every feeding guild in reef environments, and trophic variation is a central axis of diversification in this family.

Animals↗

Comparing low coverage random shotgun sequence data from Brassica oleracea and Oryza sativa genome sequence for their ability to add to the annotation of Arabidopsis thaliana.

Since the completion of the Arabidopsis thaliana genome sequence, there is an ongoing effort to annotate the genome as accurately as possible. Comparing genome sequences of related species complements the current annotation strategies by identifying genes and improving gene structure. A total of 595,321 Brassica oleracea shotgun reads were sequenced by TIGR (The Institute for Genome Research) and the collaboration of Washington University and Cold Spring Harbor. Vicogenta (a genome viewer based on GMOD and GBrowse) was created to view the current annotation and sequence alignments for Arabidopsis. Brassica reads were compared with the Arabidopsis genome and proteome databases using BLAST. Hypothetical genes and conserved unannotated regions on the short arm of chromosome 4 from Arabidopsis were experimentally verified using RT-PCR. We were able to improve the Arabidopsis annotation by identifying 25 genes that were missed, and confirming expression of 43 hypothetical genes in Arabidopsis. We were also able to detect conservation in genes whose transcription is normally suppressed due to methylation. We also examined how useful the O. sativa genome and ESTs from other species are, compared with Brassica, in improving the Arabidopsis annotation.

Amino Acid Sequence↗

Analysis of the promoter from an expanding mouse retrotransposon subfamily.

The mouse genome contains several subfamilies of the retrotransposon L1. One subfamily, TF, contains 4000-5000 full-length members and is expanding due to retrotransposition of a large number of active elements. Here we studied the TF 5' untranslated region (UTR), which contains promoter activity required for subfamily expression. Using reporter assays, we show that promoter activity is derived from TF-specific monomer sequences and is proportional to the number of monomers in the 5' UTR. These data suggest that nearly all full-length TF elements in the mouse genome are currently competent for expression. We aligned the sequences of 53 monomers to generate a consensus TF monomer and determined that most TF elements are truncated near a potential binding site for a transcription initiation factor. We also determined that much of the sequence variation among TF monomers results from transition mutations at CpG dinucleotides, suggesting that genomic TF 5' UTRs are methylated at CpGs.

Animals↗

A nested polymerase chain reaction for the detection of Borrelia burgdorferi sensu lato based on a multiple sequence analysis of the hbb gene.

A highly sensitive nested polymerase chain reaction method was designed for the detection of a wide spectrum of strains from Borrelia burgdorferi sensu lato. This technique allows the detection of as little as 3 fg of total genomic DNA extracted and purified from pure cultures of the organism, this amount corresponds to less than 10 organisms. Two sets of primers homologous to conserved spots in the coding region of the hbb gene, encoding a conserved histone-like protein, were constructed. These were based on a multiple sequence alignment of 39 strains representing all the genomic groups described in B. burgdorferi sensu lato.

Base Sequence↗

The genetics of ATP-binding cassette transporters.

The ATP-binding cassette (ABC) superfamily consists of membrane proteins that transport a wide variety of substrates across membranes. Mutations in ABC transporters cause or contribute to a number of different Mendelian disorders, including adrenoleukodystrophy, cystic fibrosis, retinal degeneration, cholesterol, and bile transport defects. In addition, the genes are involved in an increasing number of complex disorders. The proteins play essential roles in the protection of organisms from toxic metabolites and compounds in the diet and are involved in the transport of compounds across the intestine, blood-brain barrier, and the placenta. There are 48 ABC genes in the human genome divided into seven subfamilies based in gene structure, amino acid alignment, and phylogenetic analysis. These seven subfamilies are found in all other sequenced eukaryotic genomes and are of ancient origin. Further characterization of all ABC genes from humans and model organisms will lead to additional insights into normal physiology and human disease.

ATP-Binding Cassette Transporters↗

Comparative mapping of the Pseudomonas aeruginosa PAO genome with rare-cutter linking clones or two-dimensional pulsed-field gel electrophoresis protocols.

The Spe1 map of the Pseudomonas aeruginosa PAO (DSM 1707) chromosome was constructed by utilizing two-dimensional pulsed-field gel electrophoresis (PFGE) and rare-cutter linking clones. After end-labeling and fluorescence staining of macrorestriction fragments had been combined, the two-dimensional PFGE analyses of partial-total digests and reciprocal double digests were sufficient for the placement of all fragments on the genomic map. Spe1 linking fragments were isolated from BamH1, Pst1, and EcoR1 genomic libraries of P. aeruginosa PAO. After separation of the heterogeneously sized populations of Spe1-linearized and uncut circular plasmid DNAs by field inversion polyacrylamide gel electrophoresis, the gel-eluted linear DNAs were recircularized and subcloned. The 46 analyzed Spe1 linking clones recognized 16 of the 38 fragment links on the Spe1 genome map of P. aeruginosa PAO. The alignment with linking clones was consistent with that obtained from two-dimensional PFGE mapping protocols.

DNA, Bacterial↗

Seventy-five percent accuracy in protein secondary structure prediction.

In this study we present an accurate secondary structure prediction procedure by using an query and related sequences. The most novel aspect of our approach is its reliance on local pairwise alignment of the sequence to be predicted with each related sequence rather than utilization of a multiple alignment. The residue-by-residue accuracy of the method is 75% in three structural states after jack-knife tests. The gain in prediction accuracy compared with the existing techniques, which are at best 72%, is achieved by secondary structure propensities based on both local and long-range effects, utilization of similar sequence information in the form of carefully selected pairwise alignment fragments, and reliance on a large collection of known protein primary structures. The method is especially appropriate for large-scale sequence analysis of efforts such as genome characterization, where precise and significant multiple sequence alignments are not available or achievable.

Algorithms↗

RibAlign: a software tool and database for eubacterial phylogeny based on concatenated ribosomal protein subunits.

BACKGROUND: Until today, analysis of 16S ribosomal RNA (rRNA) sequences has been the de-facto gold standard for the assessment of phylogenetic relationships among prokaryotes. However, the branching order of the individual phlya is not well-resolved in 16S rRNA-based trees. In search of an improvement, new phylogenetic methods have been developed alongside with the growing availability of complete genome sequences. Unfortunately, only a few genes in prokaryotic genomes qualify as universal phylogenetic markers and almost all of them have a lower information content than the 16S rRNA gene. Therefore, emphasis has been placed on methods that are based on multiple genes or even entire genomes. The concatenation of ribosomal protein sequences is one method which has been ascribed an improved resolution. Since there is neither a comprehensive database for ribosomal protein sequences nor a tool that assists in sequence retrieval and generation of respective input files for phylogenetic reconstruction programs, RibAlign has been developed to fill this gap. RESULTS: RibAlign serves two purposes: First, it provides a fast and scalable database that has been specifically adapted to eubacterial ribosomal protein sequences and second, it provides sophisticated import and export capabilities. This includes semi-automatic extraction of ribosomal protein sequences from whole-genome GenBank and FASTA files as well as exporting aligned, concatenated and filtered sequence files that can directly be used in conjunction with the PHYLIP and MrBayes phylogenetic reconstruction programs. CONCLUSION: Up to now, phylogeny based on concatenated ribosomal protein sequences is hampered by the limited set of sequenced genomes and high computational requirements. However, hundreds of full and draft genome sequencing projects are on the way, and advances in cluster-computing and algorithms make phylogenetic reconstructions feasible even with large alignments of concatenated marker genes. RibAlign is a first step in this direction and may be particularly interesting to scientists involved in whole genome sequencing of representatives of new or sparsely studied eubacterial phyla. RibAlign is available at http://www.megx.net/ribalign.

Algorithms↗

Genome conservation between the bovine and human interleukin-8 receptor complex: improper annotation of bovine interleukin-8 receptor b identified.

Interleukin (IL)-8 and its receptors, CXCR1 and CXCR2, are key regulators of inflammation. However, knowledge of these receptors at the genomic level is limiting or absent in cattle. Therefore, our objective was to identify bovine orthologs of human CXCR1 and CXCR2. Alignment of bovine CXCR2 reference mRNA to the bovine genome revealed two regions of similarity on BTA2 approximately 20 kb apart and on opposite strands. Comparison with the human genome suggested the more centromeric region to be CXCR2 and the more telomeric region to be CXCR1 which contradicts the current annotation of the bovine CXCR2 reference mRNA. This observation was verified by sequencing RT-PCR products of specific regions within each predicted IL-8 receptor and comparing with human sequences using ClustalW. Further examination of coding and non-coding regions within the IL-8 receptor genome complex revealed that both bovine and canine CXCR1 and CXCR2 genes had more conserved sequences in common with the human genes than either mouse or rat, and may offer more suitable animal models for certain applications. This molecular information provides a stepping stone for greater understanding of the role each IL-8 receptor plays in inflammation and will enhance our ability to develop strategies against inflammatory based diseases.

Animals↗

GRS: a graphic tool for genome retrieval and segment analysis.

GRS is a graphic tool for retrieval and visualization of genome segments from partially or completely sequenced genomes. To facilitate visual identification of conserved genomic motifs, genes are color-coded according to their presumed functional roles. Aligned genes can be rapidly screened for potential homology by automatic retrieval and alignment of the corresponding protein sequences. Furthermore, the map location of any genome segment can be visually compared to the position of the same segment in other genomes or to the position of other segments within the same genome. The gene string analysis option of GRS allows the identification of genes that are identically arranged in any pairwise set of genomes. Finally, the program allows the user to create new gene table format files to enable comparisons of gene order structures in recently determined sequence data to the patterns of genes in already existing microbial and organellar databases. With the help of GRS, the genomic contexts of genes for which no identifiable homologues exist can be analyzed to provide an additional source of information for sequence annotations. We illustrate the use of GRS by analyzing the structure and distribution of phylogenetically conserved motifs in closely as well as more distantly related microbial genomes.

Computer Graphics↗

Extensive duplication and reshuffling in the Arabidopsis genome.

Systematic analysis of the Arabidopsis genome provides a basis for detailed studies of genome structure and evolution. Members of multigene families were mapped, and random sequence alignment was used to identify regions of extended similarity in the Arabidopsis genome. Detailed analysis showed that the number, order, and orientation of genes were conserved over large regions of the genome, revealing extensive duplication covering the majority of the known genomic sequence. Fine mapping analysis showed much rearrangement, resulting in a patchwork of duplicated regions that indicated deletion, insertion, tandem duplication, inversion, and reciprocal translocation. The implications of these observations for evolution of the Arabidopsis genome as well as their usefulness for analysis and annotation of the genomic sequence and in comparative genomics are discussed.

Arabidopsis↗

The web server of IBM's Bioinformatics and Pattern Discovery group.

We herein present and discuss the services and content which are available on the web server of IBM's Bioinformatics and Pattern Discovery group. The server is operational around the clock and provides access to a variety of methods that have been published by the group's members and collaborators. The available tools correspond to applications ranging from the discovery of patterns in streams of events and the computation of multiple sequence alignments, to the discovery of genes in nucleic acid sequences and the interactive annotation of amino acid sequences. Additionally, annotations for more than 70 archaeal, bacterial, eukaryotic and viral genomes are available on-line and can be searched interactively. The tools and code bundles can be accessed beginning at http://cbcsrv.watson.ibm.com/Tspd.html whereas the genomics annotations are available at http://cbcsrv.watson.ibm.com/Annotations/.

Computational Biology↗

Evidence for diversifying selection at the pyoverdine locus of Pseudomonas aeruginosa.

Pyoverdine is the primary siderophore of the gram-negative bacterium Pseudomonas aeruginosa. The pyoverdine region was recently identified as the most divergent locus alignable between strains in the P. aeruginosa genome. Here we report the nucleotide sequence and analysis of more than 50 kb in the pyoverdine region from nine strains of P. aeruginosa. There are three divergent sequence types in the pyoverdine region, which correspond to the three structural types of pyoverdine. The pyoverdine outer membrane receptor fpvA may be driving diversity at the locus: it is the most divergent alignable gene in the region, is the only gene that showed substantial intratype variation that did not appear to be generated by recombination, and shows evidence of positive selection. The hypothetical membrane protein PA2403 also shows evidence of positive selection; residues on one side of the membrane after protein folding are under positive selection. R', previously identified as a type IV strain, is clearly derived from a type III strain via a 3.4-kb deletion which removes one amino acid from the pyoverdine side chain peptide. This deletion represents a natural modification of the product of a nonribosomal peptide synthetase enzyme, whose consequences are predictive from the DNA sequence. There is also linkage disequilibrium between the pyoverdine region and pvdY, a pyoverdine gene separated by 30 kb from the pyoverdine region. The pyoverdine region shows evidence of horizontal transfer; we propose that some alleles in the region were introduced from other soil bacteria and have been subsequently maintained by diversifying selection.

Bacterial Outer Membrane Proteins↗

Automatic assessment of alignment quality.

Multiple sequence alignments play a central role in the annotation of novel genomes. Given the biological and computational complexity of this task, the automatic generation of high-quality alignments remains challenging. Since multiple alignments are usually employed at the very start of data analysis pipelines, it is crucial to ensure high alignment quality. We describe a simple, yet elegant, solution to assess the biological accuracy of alignments automatically. Our approach is based on the comparison of several alignments of the same sequences. We introduce two functions to compare alignments: the average overlap score and the multiple overlap score. The former identifies difficult alignment cases by expressing the similarity among several alignments, while the latter estimates the biological correctness of individual alignments. We implemented both functions in the MUMSA program and demonstrate the overall robustness and accuracy of both functions on three large benchmark sets.

Algorithms↗

Analysis of bovine mammary gland EST and functional annotation of the Bos taurus gene index.

Functional genomic studies of the mammary gland require an appropriate collection of cDNA sequences to assess gene expression patterns from the different developmental and operational states of underlying cell types. To better capture the range of gene expression, a normalized cDNA library was constructed from pooled bovine mammary tissues, and 23,202 expressed sequence tags (EST) were produced and deposited into GenBank. Assembly of these EST with sequences in the Bos taurus Gene Index (BtGI) helped to form 5751 of the current 23,883 tentative consensus (TC) sequences. The majority (87%) of these 5751 assemblies contained only one to three mammary-derived EST. In contrast, 18% of the mammary EST assembled with TC sequences corresponding to 12 genes. These results suggest library normalization was only partially effective, because the reduction in EST for genes abundantly transcribed during lactation could be attributed to pooling. For better assessment of novel content in the mammary library and to add to existing annotation of all bovine sequence elements, gene ontology assignments, and comparative sequence analyses against human genome sequence, human and rodent gene indices, and an index of orthologous alignments of genes across eukaryotes (TOGA) were performed, and results were added to existing BtGI annotation. Over 35,000 of the bovine elements significantly matched human genome sequence, and the positions of some alignments (3%) were unique relative to those using human expressed sequences. Because 3445 TC sequences had no significant match with any data set, mammary-derived cDNA clones representing 23 of these elements were analyzed further for expression and novelty. Only one clone met criteria suggesting the corresponding gene was a divergent ortholog or expressed sequence unique to cattle. These results demonstrate that bovine sequence expression data serve as a resource for characterizing mammalian transcriptomes and identifying those genes potentially unique to ruminants.

Animals↗

Rat growth hormone gene introns stimulate nucleosome alignment in vitro and in transgenic mice.

Average hepatic expression (mRNA per cell per gene) of a metallothionein-rat growth hormone (rGH) gene with its natural introns was about 15-fold higher than an intronless version when tested in transgenic mice. We examined the idea that intron removal leads to an alteration in chromatin structure that might be responsible for this effect. Using an in vitro chromatin assembly system, we observed that nucleosomes were aligned in a characteristic ordered array over the gene and promoter when all introns were present. Linker histones were necessary for this alignment to occur. In contrast, nucleosome alignment was perturbed in constructs lacking some or all of the introns. A similar disruption of nucleosome alignment was observed when comparing chromatin from livers of transgenic mice carrying rGH transgenes with or without introns. In vitro, sequences at the 3' end of the rGH gene position nucleosomes and facilitate nucleosome alignment upstream; however, nucleosome alignment does not occur on the approximately 3 kb of downstream flanking rat sequence. These observations suggest that signals present in genomic rGH DNA may serve to establish appropriate nucleosome alignment during development and, possibly, to restore nucleosome alignment to the transcribed region after disruption incurred by the passage of an RNA polymerase molecule, thereby facilitating subsequent rounds of transcription.

Animals↗

Following tetraploidy in an Arabidopsis ancestor, genes were removed preferentially from one homeolog leaving clusters enriched in dose-sensitive genes.

Approximately 90% of Arabidopsis' unique gene content is found in syntenic blocks that were formed during the most recent whole-genome duplication. Within these blocks, 28.6% of the genes have a retained pair; the remaining genes have been lost from one of the homeologs. We create a minimized genome by condensing local duplications to one gene, removing transposons, and including only genes within blocks defined by retained pairs. We use a moving average of retained and non-retained genes to find clusters of retention and then identify the types of genes that appear in clusters at frequencies above expectations. Significant clusters of retention exist for almost all chromosomal segments. Detailed alignments show that, for 85% of the genome, one homeolog was preferentially (1.6x) targeted for fractionation. This homeolog fractionation bias suggests an epigenetic mechanism. We find that islands of retention contain "connected genes," those genes predicted-by the gene balance hypothesis-to be resistant to removal because the products they encode interact with other products in a dose-sensitive manner, creating a web of dependency. Gene families that are overrepresented in clusters include those encoding components of the proteasome/protein modification complexes, signal transduction machinery, ribosomes, and transcription factor complexes. Gene pair fractionation following polyploidy or segmental duplication leaves a genome enriched for "connected" genes. These clusters of duplicate genes may help explain the evolutionary origin of coregulated chromosomal regions and new functional modules.

Arabidopsis↗