Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

[Preliminary study of the gene structure of human glycosylphosphatidylinositol specific phospholipase D].

To explore the cDNA and its genomic gene structure of human glycosylphosphatidylinositol specific phospholipase D (GPI-PLD), the stromal cells from human bone marrow were cultured, and the GPI-PLD cDNA was successfully cloned from stromal cells (GenBank Accession AY007546). After analyzing this cDNA, we found that it codes the signal peptide as well as the mature peptide (817 AA) of GPI-PLD. Compared with the cDNA cloned from human pancreas and liver, the homology of the cDNA is 99% and 95% respectively. After searching GenBank, we aligned the genomic gene of the GPI-PLD by DNATools software, and found that the GPI-PLD genomic gene was located in the human 6p22.1-22.3 and contained 25 exons, TATA box, CAAT box, enhancer and the sequence binding homo domain of Pit-1/GHF-1.

Bone Marrow Cells↗

Quantitative DNA methylation analysis based on four-dye trace data from direct sequencing of PCR amplificates.

MOTIVATION: Methylation of cytosines in DNA plays an important role in the regulation of gene expression, and the analysis of methylation patterns is fundamental for the understanding of cell differentiation, aging processes, diseases and cancer development. Such analysis has been limited, because technologies for detailed and efficient high-throughput studies have not been available. We have developed a novel quantitative methylation analysis algorithm and workflow based on direct DNA sequencing of PCR products from bisulfite-treated DNA with high-throughput sequencing machines. This technology is a prerequisite for success of the Human Epigenome Project, the first large genome-wide sequencing study for DNA methylation in many different tissues. Methylation in tissue samples which are compositions of different cells is a quantitative information represented by cytosine/thymine proportions after bisulfite conversion of unmethylated cytosines to uracil and PCR. Calculation of quantitative methylation information from base proportions represented by different dye signals in four-dye sequencing trace files needs a specific algorithm handling imbalanced and overscaled signals, incomplete conversion, quality problems and basecaller artifacts. RESULTS: The algorithm we developed has several key properties: it analyzes trace files from PCR products of bisulfite-treated DNA sequenced directly on ABI machines; it yields quantitative methylation measurements for individual cytosine positions after alignment with genomic reference sequences, signal normalization and estimation of effectiveness of bisulfite treatment; it works in a fully automated pipeline including data quality monitoring; it is efficient and avoids the usual cost of multiple sequencing runs on subclones to estimate DNA methylation. The power of our new algorithm is demonstrated with data from two test systems based on mixtures with known base compositions and defined methylation. In addition, the applicability is proven by identifying CpGs that are differentially methylated in real tissue samples.

Algorithms↗

M-GCAT: interactively and efficiently constructing large-scale multiple genome comparison frameworks in closely related species.

BACKGROUND: Due to recent advances in whole genome shotgun sequencing and assembly technologies, the financial cost of decoding an organism's DNA has been drastically reduced, resulting in a recent explosion of genomic sequencing projects. This increase in related genomic data will allow for in depth studies of evolution in closely related species through multiple whole genome comparisons. RESULTS: To facilitate such comparisons, we present an interactive multiple genome comparison and alignment tool, M-GCAT, that can efficiently construct multiple genome comparison frameworks in closely related species. M-GCAT is able to compare and identify highly conserved regions in up to 20 closely related bacterial species in minutes on a standard computer, and as many as 90 (containing 75 cloned genomes from a set of 15 published enterobacterial genomes) in an hour. M-GCAT also incorporates a novel comparative genomics data visualization interface allowing the user to globally and locally examine and inspect the conserved regions and gene annotations. CONCLUSION: M-GCAT is an interactive comparative genomics tool well suited for quickly generating multiple genome comparisons frameworks and alignments among closely related species. M-GCAT is freely available for download for academic and non-commercial use at: http://alggen.lsi.upc.es/recerca/align/mgcat/intro-mgcat.html.

Algorithms↗

Comparative analysis of 87,000 expressed sequence tags from the fumonisin-producing fungus Fusarium verticillioides.

Fusarium verticillioides (teleomorph Gibberella moniliformis) is a pathogen of maize worldwide and produces fumonisins, a family of mycotoxins that have been associated with several animal diseases as well as cancer in humans. In this study, we sought to identify fungal genes that affect fumonisin production and/or the plant-fungal interaction. We generated over 87,000 expressed sequence tags from nine different cDNA libraries that correspond to 11,119 unique sequences and are estimated to represent 80% of the genomic complement of genes. A comparative analysis of the libraries showed that all 15 genes in the fumonisin gene cluster were differentially expressed. In addition, nine candidate fumonisin regulatory genes and a number of genes that may play a role in plant-fungal interaction were identified. Analysis of over 700 FUM gene transcripts from five different libraries provided evidence for transcripts with unspliced introns and spliced introns with alternative 3' splice sites. The abundance of the alternative splice forms and the frequency with which they were found for genes involved in the biosynthesis of a single family of metabolites as well as their differential expression suggest they may have a biological function. Finally, analysis of an EST that aligns to genomic sequence between FUM12 and FUM13 provided evidence for a previously unidentified gene (FUM20) in the FUM gene cluster.

Amino Acid Sequence↗

Splice variation in mouse full-length cDNAs identified by mapping to the mouse genome.

We mapped the collection of The Institute of Physical and Chemical Research (Japan) (RIKEN) 21,076 full-length mouse cDNA clone sequences and the mouse RefSeq sequences to the recently completed draft of the mouse genome. Using this mapping, we identified 3674 mouse genes with multiple transcripts, of which 1098 have splice variants. All but 532 of 21,076 clones (97.5%) mapped to the genome assembly. Alignments of cDNA clone sequences with proteins show that much of the detected splice variation alters coding regions and affects the translated protein. We developed novel analytical techniques to classify observed splice variation and to assess the relation between splice variation and alternative transcription. This analysis indicates that an alternative choice of transcription start or polyadenylation signal frequently induces splice variation.

Alternative Splicing↗

A jumping profile Hidden Markov Model and applications to recombination sites in HIV and HCV genomes.

BACKGROUND: Jumping alignments have recently been proposed as a strategy to search a given multiple sequence alignment A against a database. Instead of comparing a database sequence S to the multiple alignment or profile as a whole, S is compared and aligned to individual sequences from A. Within this alignment, S can jump between different sequences from A, so different parts of S can be aligned to different sequences from the input multiple alignment. This approach is particularly useful for dealing with recombination events. RESULTS: We developed a jumping profile Hidden Markov Model (jpHMM), a probabilistic generalization of the jumping-alignment approach. Given a partition of the aligned input sequence family into known sequence subtypes, our model can jump between states corresponding to these different subtypes, depending on which subtype is locally most similar to a database sequence. Jumps between different subtypes are indicative of intersubtype recombinations. We applied our method to a large set of genome sequences from human immunodeficiency virus (HIV) and hepatitis C virus (HCV) as well as to simulated recombined genome sequences. CONCLUSION: Our results demonstrate that jumps in our jumping profile HMM often correspond to recombination breakpoints; our approach can therefore be used to detect recombinations in genomic sequences. The recombination breakpoints identified by jpHMM were found to be significantly more accurate than breakpoints defined by traditional methods based on comparing single representative sequences.

Algorithms↗

Complete genome sequence of the broad-host-range vibriophage KVP40: comparative genomics of a T4-related bacteriophage.

The complete genome sequence of the T4-like, broad-host-range vibriophage KVP40 has been determined. The genome sequence is 244,835 bp, with an overall G+C content of 42.6%. It encodes 386 putative protein-encoding open reading frames (CDSs), 30 tRNAs, 33 T4-like late promoters, and 57 potential rho-independent terminators. Overall, 92.1% of the KVP40 genome is coding, with an average CDS size of 587 bp. While 65% of the CDSs were unique to KVP40 and had no known function, the genome sequence and organization show specific regions of extensive conservation with phage T4. At least 99 KVP40 CDSs have homologs in the T4 genome (Blast alignments of 45 to 68% amino acid similarity). The shared CDSs represent 36% of all T4 CDSs but only 26% of those from KVP40. There is extensive representation of the DNA replication, recombination, and repair enzymes as well as the viral capsid and tail structural genes. KVP40 lacks several T4 enzymes involved in host DNA degradation, appears not to synthesize the modified cytosine (hydroxymethyl glucose) present in T-even phages, and lacks group I introns. KVP40 likely utilizes the T4-type sigma-55 late transcription apparatus, but features of early- or middle-mode transcription were not identified. There are 26 CDSs that have no viral homolog, and many did not necessarily originate from Vibrio spp., suggesting an even broader host range for KVP40. From these latter CDSs, an NAD salvage pathway was inferred that appears to be unique among bacteriophages. Features of the KVP40 genome that distinguish it from T4 are presented, as well as those, such as the replication and virion gene clusters, that are substantially conserved.

Bacteriophage T4↗

Phylogenomic subsampling and upsampling for efficient evolutionary analyses of big data.

Long runtimes, high memory demands, and reliance on high-performance computing impede phylogenomic analyses. We review a scalable phylogenomic subsampling with upsampling (PSU) framework, in which small subsamples of sites from a concatenated alignment are expanded by upsampling before inference, and the resulting analyses are then aggregated to obtain evolutionary estimates. PSU harnesses the fact that the computational cost of maximum likelihood analysis is strongly influenced by the number of distinct site patterns in the concatenated alignment, whereas statistical power depends primarily on the amount of evolutionary information represented by the total number of sites and substitutions. By reducing the former while restoring the latter through upsampling, PSU can approximate many full-data analyses at substantially lower computational cost. Analysis of simulated and empirical datasets shows that PSU can accurately estimate bootstrap support values, select the optimal substitution model, test evolutionary hypotheses, and infer branch lengths, divergence times, and associated uncertainty measures, while reducing runtime and memory requirements by orders of magnitude. PSU also provides distributions of inferred clade support across independent subsamples, enabling detection of conflicting phylogenetic signals that may remain hidden in conventional bootstrap analysis. Automated tuning of subsample size, the number of subsamples, and the number of upsampling replicates make PSU practical across diverse datasets. We suggest that PSU is a general strategy for scalable phylogenomic inference using a broad range of statistical methods. By enabling analyses of genome-scale alignments on commodity hardware, PSU broadens research access and reduces environmental and infrastructural costs of big-data phylogenomics.

confidence limits↗

Is there a phylogenetic signal in prokaryote proteins?

Using the sequence information from nine completely sequenced bacterial genomes, we extract 32 protein families that are thought to contain orthologous proteins from each genome. The alignments of these 32 families are used to construct a phylogeny with the neighbor-joining algorithm. This tree has several topological features that are different from the conventional phylogeny, yet it is highly reliable according to its bootstrap values. Upon closer study of the individual families used, it is clear that the strong phylogenetic signal comes from three families, at least two of which are good candidates for horizontal transfer. The tree from the remaining 29 families consists almost entirely of noise at the level of bacterial phylum divisions, indicating that, even with large amounts of data, it may not be possible to reconstruct the prokaryote phylogeny using standard sequence-based methods.

Arginine-tRNA Ligase↗

Functional classification, genomic organization, putatively cis-acting regulatory elements, and relationship to quantitative trait loci, of sorghum genes with rhizome-enriched expression.

Rhizomes are organs of fundamental importance to plant competitiveness and invasiveness. We have identified genes expressed at substantially higher levels in rhizomes than other plant parts, and explored their functional categorization, genomic organization, regulatory motifs, and association with quantitative trait loci (QTLs) conferring rhizomatousness. The finding that genes with rhizome-enriched expression are distributed across a wide range of functional categories suggests some degree of specialization of individual members of many gene families in rhizomatous plants. A disproportionate share of genes with rhizome-enriched expression was implicated in secondary and hormone metabolism, and abiotic stimuli and development. A high frequency of unknown-function genes reflects our still limited knowledge of this plant organ. A putative oligosaccharyl transferase showed the highest degree of rhizome-specific expression, with several transcriptional or regulatory protein complex factors also showing high (but lesser) degrees of specificity. Inferred by the upstream sequences of their putative rice (Oryza sativa) homologs, sorghum (Sorghum bicolor) genes that were relatively highly expressed in rhizome tip tissues were enriched for cis-element motifs, including the pyrimidine box, TATCCA box, and CAREs box, implicating the gibberellins in regulation of many rhizome-specific genes. From cDNA clones showing rhizome-enriched expression, expressed sequence tags forming 455 contigs were plotted on the rice genome and aligned to QTL likelihood intervals for ratooning and rhizomatous traits in rice and sorghum. Highly expressed rhizome genes were somewhat enriched in QTL likelihood intervals for rhizomatousness or ratooning, with specific candidates including some of the most rhizome-specific genes. Some rhizomatousness and ratooning QTLs were shown to be potentially related to one another as a result of ancient duplication, suggesting long-term functional conservation of the underlying genes. Insight into genes and pathways that influence rhizome growth set the stage for genetic and/or exogenous manipulation of rhizomatousness, and for further dissection of the molecular evolution of rhizomatousness.

Chromosome Mapping↗

Phylogeny based discovery of regulatory elements.

BACKGROUND: Algorithms that locate evolutionarily conserved sequences have become powerful tools for finding functional DNA elements, including transcription factor binding sites; however, most methods do not take advantage of an explicit model for the constrained evolution of functional DNA sequences. RESULTS: We developed a probabilistic framework that combines an HKY85 model, which assigns probabilities to different base substitutions between species, and weight matrix models of transcription factor binding sites, which describe the probabilities of observing particular nucleotides at specific positions in the binding site. The method incorporates the phylogenies of the species under consideration and takes into account the position specific variation of transcription factor binding sites. Using our framework we assessed the suitability of alignments of genomic sequences from commonly used species as substrates for comparative genomic approaches to regulatory motif finding. We then applied this technique to Saccharomyces cerevisiae and related species by examining all possible six base pair DNA sequences (hexamers) and identifying sequences that are conserved in a significant number of promoters. By combining similar conserved hexamers we reconstructed known cis-regulatory motifs and made predictions of previously unidentified motifs. We tested one prediction experimentally, finding it to be a regulatory element involved in the transcriptional response to glucose. CONCLUSION: The experimental validation of a regulatory element prediction missed by other large-scale motif finding studies demonstrates that our approach is a useful addition to the current suite of tools for finding regulatory motifs.

Algorithms↗

Elucidation of factors responsible for enhanced thermal stability of proteins: a structural genomics based study.

Understanding the molecular basis for the enhanced stability of proteins from thermophiles has been hindered by a lack of structural data for homologous pairs of proteins from thermophiles and mesophiles. To overcome this difficulty, complete genome sequences from 9 thermophilic and 21 mesophilic bacterial genomes were aligned with protein sequences with known structures from the protein data bank. Sequences with high homology to proteins with known structures were chosen for further analysis. High quality models of these chosen sequences were obtained using homology modeling. The current study is based on a data set of models of 900 mesophilic and 300 thermophilic protein single chains and also includes 178 templates of known structure. Structural comparisons of models of homologous proteins allowed several factors responsible for enhanced thermostability to be identified. Several statistically significant, specific amino acid substitutions that occur going from mesophiles to thermophiles are identified. Most of these are at solvent-exposed sites. Salt bridges occur significantly more often in thermophiles. The additional salt bridges in thermophiles are almost exclusively in solvent-exposed regions, and 35% are in the same element of secondary structure. Helices in thermophiles are stabilized by intrahelical salt bridges and by an increase in negative charge at the N-terminus. There is an approximate decrease of 1% in the overall loop content and a corresponding increase in helical content in thermophiles. Previously overlooked cation-pi interactions, estimated to be twice as strong as ion-pairs, are significantly enriched in thermophiles. At buried sites, statistically significant hydrophobic amino acid substitutions are typically consistent with decreased side chain conformational entropy.

Archaeal Proteins↗

Complete sequence determination of the mouse and human CTLA4 gene loci: cross-species DNA sequence similarity beyond exon borders.

CTLA4 (CD152), a receptor for the B7 costimulatory molecules (CD80 and CD86), is considered a fundamental regulator of T-cell activation. In this paper, we present the complete primary structure of the mouse and human CTLA4 gene loci. Sequence comparison between the mouse and the human CTLA4 gene loci revealed a high degree of sequence conservation both for homologous noncoding regions (65-78% identity) and for coding regions (72-98% identity), with an overall score of 71% over the entire length of the two genes. Of the CTLA4 genomic regions aligned, five simple repetitive elements were found in the mouse locus, whereas two simple repetitive sequences were localized on the human locus. RNA blot analysis of mouse and human primary tissues indicated that both CTLA4 and T-cell receptor transcripts were found in most organs with generally higher levels in lymphoid tissues. The conservation of CTLA4 gene patterning raises the possibility that constrained gene evolution of CTLA4 may be linked to conserved transcriptional control of this locus.

Abatacept↗

A new zinc ribbon gene (ZNRD1) is cloned from the human MHC class I region.

Eleven unique cDNA fragments were identified from YAC B30H3, which spans 330 kb in the human major histocompatibility complex class I region. One fragment (CAT80) was mapped 80 kb telomeric to the HLA-A locus. Using this cDNA fragment as probe, Northern analysis reveals a ubiquitously expressed transcript of about 850 nt in all 16 tissues tested. Based on the cDNA fragment sequence, a full-length cDNA of 858 bp that contains an open reading frame of 378 bp was cloned. Within the putative polypeptide of 126 amino acids, two zinc-ribbon domains were identified: Cx2Cx15Cx2C at the N-terminal and Cx2Cx24Cx2C at the C-terminal. The C-terminal domain is well conserved throughout evolution, including archaea, yeast, Drosophila, nematodes, amphibians, and mammals. The conserved amino acid sequence, CxRCx6Yx3QxRSADEx2TxFxCx2C, is highly homologous to the yeast RNA polymerase A subunit 9 and transcription-associated proteins. Alignment with genomic DNA demonstrates that this gene spans 3.6 kb and consists of four exons and three introns. Cross-species Northern analysis reveals a mouse homolog of a similar size and with an expression profile similar to those of the human gene. We have named this gene ZNRD1 for zinc ribbon domain-containing 1 protein.

Amino Acid Sequence↗

cDNA cloning of a novel human gene NAKAP95, neighbor of A-kinase anchoring protein 95 (AKAP95) on chromosome 19p13.11-p13.12 region.

A-kinase anchoring protein 95 (AKAP95) is a nuclear protein which binds to the regulatory subunit (RII) of cyclic adenosine monophosphate (cAMP)-dependent protein kinase (PKA) and to DNA. A novel nuclear human gene which shares sequence homology with the human AKAP95 gene was identified by a nuclear transportation trap method. By polymerase chain reaction (PCR)-based analysis with both a human/rodent monochromosomal hybrid cell panel and a radiation hybrid panel, the gene was mapped to the chromosome 19p13.11-p13.12 region between markers WI-4669 and CHLC.GATA27C12. Furthermore, alignment with genomic sequences revealed that the gene and human AKAP95 resided tandemly only approximately 250 bp apart from each other. We designated this gene as neighbor of AKAP95 (NAKAP95). The exon-intron structure of NAKAP95 and AKAP95 was conserved, indicating that they may have evolved by gene duplication. The predicted protein product of the NAKAP95 gene consists of 646 amino acid residues, and NAKAP95 and AKAP95 had an overall 40% similarity, both having a potential nuclear localizing signal and two C2H2 type zinc finger motifs. The putative RII binding motif in AKAP95 was not conserved in NAKAP95. A reverse transcription coupled (RT)-PCR experiment revealed that the NAKAP95 gene was transcribed ubiquitously in various human tissues.

Amino Acid Sequence↗

A novel member of the serpin superfamily is encoded on a circular plasmid-like DNA species isolated from rabbit cells.

A novel member of the serpin family of serine protease inhibitors is presented. A plasmid-like DNA was isolated from rabbit cells by its homology to the genome of Shope fibroma virus (SFV), a tumorigenic poxvirus of rabbits, and was shown elsewhere to encode a serpin-like protein [(1986) Mol. Cell. Biol. 6, 265-276]. Although significant DNA homology exists between the rabbit plasmid serpin open reading frame and the SFV terminal inverted repeat DNA there is no intact serpin counterpart encoded by this region of the SFV genome. The alignment of the novel plasmid-borne polypeptide with the serpin family of proteins confirms its status within this group.

Animals↗

Tunicate muscle actin genes. Structure and organization as a gene cluster.

We have isolated and determined the complete nucleotide sequences of two genes, HrMA4a and HrMA2, which encode the same muscle actin protein of the tunicate Halocynthia roretzi. HrMA4a and HrMA2 contain three exons, and the genes have intron-exon splice junctions at the same positions. The 5' flanking region of HrMA4a gene contains several potential regulatory elements. A TATA box is located at -30 and a CArG box found in regulatory region of vertebrate muscle-specific genes is located at -116. Seven E-box consensus sequences (CANNTG) known as binding sites for vertebrate myogenic determination factors are found within a 500 base-pair portion of the 5' flanking region of HrMA4a gene. HrMA4a and HrMA2 are separated by 1600 bases in genomic DNA and transcribed in the same direction. In addition to these genes, we have identified three other actin genes encoding muscle-type actins. All five actin genes are located in a 30 x 10(3) base-pair region of the genome and aligned in the same direction. This is the first report of a cluster of "vertebrate-type" muscle actin genes. The consensus sequences of 5' flanking region are conserved among these five genes, suggesting that the expression of the genes is controlled coordinately. This may be advantageous for the accumulation of considerable amounts of actin proteins in rapidly developing embryos of this animal.

Actins↗

Sequence divergence yet conserved physical characteristics among the E4 proteins of cutaneous human papillomaviruses.

Human papillomavirus (HPV) types 1, 2, and 4 together comprise the major cause of cutaneous papillomas in the general population. We have aligned the genomes of these three viruses by partial sequence analysis, and have sequenced the E4 open reading frames (ORFs) of HPV 2 and HPV 4. After expression as beta-gal fusion proteins in bacteria, antibodies raised to the putative E4 gene-products of both virus types were used to identify the native E4 proteins in naturally occurring tumors. At the primary amino acid sequence level, the E4 protein of HPV 2 was found to be most homologous with those of HPV 6 and 11 and was not closely related to those of HPV 1 or 4. Although the E4 ORF represents a region of weak homology amongst papillomaviruses, the E4 encoded proteins showed significant conservation in their physical characteristics. Like those of HPV 1, the E4 proteins of both HPV 2 and HPV 4 were found to be composed of a major low-molecular-weight doublet (16.5/18K for HPV 2, 20/21K for HPV 4, c.f. 16/17K for HPV 1) along with minor high-molecular-weight species, which probably represent dimers of the smaller proteins, (33K for HPV 2, 40K for HPV 4, c.f. 32/34K for HPV 1). The E4 products of all three virus types were multiply charged, and exhibited a characteristic migration pattern following alkaline urea gel electrophoresis. Although the levels of E4 expression in tumors induced by the different virus types was very different, this was found to correlate closely with the level of virus production characteristic of each virus type. In all three cases, E4 proteins were found to be primarily cytoplasmic, and to be associated with the distinctive cytoplasmic inclusion granules characteristic of each virus type. The poor sequence conservation between the E4 protein of HPVs 1, 2, and 4, taken alongside the ability of these viruses to infect similar histological sites, suggests that E4 may not be involved in determining tissue specificity. Our results suggest conserved physical characteristics (acidic, multiply charged, ability to form dimers) and similar site of expression may be the important factors for E4 function.

Amino Acid Sequence↗