Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

inGeno--an integrated genome and ortholog viewer for improved genome to genome comparisons.

BACKGROUND: Systematic genome comparisons are an important tool to reveal gene functions, pathogenic features, metabolic pathways and genome evolution in the era of post-genomics. Furthermore, such comparisons provide important clues for vaccines and drug development. Existing genome comparison software often lacks accurate information on orthologs, the function of similar genes identified and genome-wide reports and lists on specific functions. All these features and further analyses are provided here in the context of a modular software tool "inGeno" written in Java with Biojava subroutines. RESULTS: InGeno provides a user-friendly interactive visualization platform for sequence comparisons (comprehensive reciprocal protein--protein comparisons) between complete genome sequences and all associated annotations and features. The comparison data can be acquired from several different sequence analysis programs in flexible formats. Automatic dot-plot analysis includes output reduction, filtering, ortholog testing and linear regression, followed by smart clustering (local collinear blocks; LCBs) to reveal similar genome regions. Further, the system provides genome alignment and visualization editor, collinear relationships and strain-specific islands. Specific annotations and functions are parsed, recognized, clustered, logically concatenated and visualized and summarized in reports. CONCLUSION: As shown in this study, inGeno can be applied to study and compare in particular prokaryotic genomes against each other (gram positive and negative as well as close and more distantly related species) and has been proven to be sensitive and accurate. This modular software is user-friendly and easily accommodates new routines to meet specific user-defined requirements.

Base Sequence↗

Gene-level complexity explains genome-wide variation in the distribution of fitness effects.

The distribution of fitness effects (DFE)-describing how harmful, neutral, or beneficial new mutations are-is central to understanding how populations evolve. Although the DFE varies across genomes and species, it remains unclear which aspects of genomic organization drive this variation. Here, we inferred gene-level selective constraints across the genomes of Mus musculus castaneus, Drosophila melanogaster and Saccharomyces cerevisiae using a combination of population genetics and machine learning trained on diverse gene features. Many gene features were predictive of selective constraint, with conservation, gene structure, and expression being the most informative. These selective constraints delineated gene classes with distinct DFEs. Genes with higher connectivity and expression-features reflecting how many traits a gene influences-experienced stronger and less dispersed deleterious effects with increasing selective constraint. Between species, the rate of adaptation decreased with increasing organismal complexity, whereas across the genome it did not decrease monotonically with selective constraint, but tended to be higher at intermediate levels. While between-species comparisons of DFE parameters were less consistent with predictions of Fisher's geometric model (FGM) based on organismal complexity, variation in DFE parameters across the genome aligned more closely with FGM when complexity was considered at the gene level. Our results suggest that gene-level complexity, captured by genomic feature proxies, provides a more informative definition of complexity for DFE variation than organism-level labels, and highlight the value of using gene features collectively to link genomic architecture, fitness landscapes, and patterns of molecular evolution.

Animals↗

Identification of alternatively spliced mRNA variants related to cancers by genome-wide ESTs alignment.

Several databases have been published to predict alternative splicing of mRNAs by analysing the exon linkage relationship by alignment of expressed sequence tags (ESTs) to the genome sequence; however, little effort has been made to investigate the relationship between cancers and alternative splicing. We developed a program, Alternative Splicing Assembler (ASA), to look for splicing variants of human gene transcripts by genome-wide ESTs alignment. Using ASA, we constructed the biosino alternative splicing database (BASD), which predicted splicing variants for reference sequences from the reference sequence database (RefSeq) and presented them in both graph and text formats. EST clusters that differ from the reference sequences in at least one splicing site were counted as splicing variants. Of 4322 genes screened, 3490 (81%) were observed with at least one alternative splicing variants. To discover the variants associated with cancers, tissue sources of EST sequences were extracted from the UniLib database and ESTs from the same tissue type were counted. These were regarded as the indicators for gene expression level. Using Fisher's exact test, alternative splicing variants, of which EST counts were significantly different between cancer tissues and their counterpart normal tissues, were identified. It was predicted that 2149 variants, or 383 variants after Bonferroni correction, of 26 812 variants were likely tumor-associated. By reverse transcription-PCR, 11 of 13 novel alternative splicing variants and eight of nine variants' tissue specificity were confirmed in hepatocellular carcinoma and in lung cancer. The possible involvement of alternative splicing in cancer is discussed.

Alternative Splicing↗

Improvements in the HbVar database of human hemoglobin variants and thalassemia mutations for population and sequence variation studies.

HbVar (http://globin.cse.psu.edu/globin/hbvar/) is a relational database developed by a multi-center academic effort to provide up-to-date and high quality information on the genomic sequence changes leading to hemoglobin variants and all types of thalassemia and hemoglobinopathies. Extensive information is recorded for each variant and mutation, including sequence alterations, biochemical and hematological effects, associated pathology, ethnic occurrence and references. In addition to the regular updates to entries, we report two significant advances: (i) The frequencies for a large number of mutations causing beta-thalassemia in at-risk populations have been extracted from the published literature and made available for the user to query upon. (ii) HbVar has been linked with the GALA (Genome Alignment and Annotation database, available at http://globin.cse.psu.edu/gala/) so that users can combine information on hemoglobin variants and thalassemia mutations with a wide spectrum of genomic data. It also expands the capacity to view and analyze the data, using tools within GALA and the University of California at Santa Cruz (UCSC) Genome Browser.

Databases, Genetic↗

Transcriptome and genome conservation of alternative splicing events in humans and mice.

Combining mRNA and EST data in splicing graphs with whole genome alignments, we discover alternative splicing events that are conserved in both human and mouse transcriptomes. 1,964 of 19,156 (10%) loci examined contain one or more such alternative splicing events, with 2,698 total events. These events represent a lower bound on the amount of alternative splicing in the human genome. Also, as these alternative splicing events are conserved between the human and mouse transcriptomes they should be enriched for functionally significant alternative splicing events, free from much of the noise found in the EST libraries. Further classification of these alternative splicing events reveals that 1,037 (38.4%) are due to exon skipping, 497 (18.4%) are due to alternative 3' splice sites, 214 (7.9%) are due to alternative 5' splice sites, 75 (2.8%) are due to intron retention and the other 875 (32.4%) are due to other, more complicated, alternative splicing events. In addition, genomic sequences nearby these alternative splicing events display increased sequence conservation. Both the alternatively spliced exons and the proximal intron show increased levels of genomic conservation relative to constitutively spliced exons. For exon skipping events both intron regions flanking the exon are conserved while for alternative 5' and 3' splicing events the conservation is greater near the alternative splice site.

Algorithms↗

Mining ORESTES no-match database: can we still contribute to cancer transcriptome?

The Human Cancer Genome Project generated about 1 million expressed sequence tags by the ORESTES method, principally with the aim of obtaining data from cancer. Of this total, 341,680 showed no similarity with sequences in the public transcript databases, referred to as "no-match". Some of them represent low abundance or difficult to detect human transcripts, but part of these sequences represent genomic contamination or immature mRNA. We performed a bioinformatics pipeline to determine the novelty of ORESTES "no-match" datasets from prostate or breast tissues. We started with 14,908 clusters mapped on the human genome. A total of 2226 clusters originating from more than two libraries or singletons with gaps upon genome alignment were selected. Ninety-four clusters with canonical splice sites representing the most stringent criteria to be considered a gene were subjected to manual inspection regarding genomic hits. Of the manually inspected clusters, 49.6% contained new sequences where 42.2% were probable low-expression alternative forms of the characterized genes and 7.4% unpredicted genes. RT-PCR followed by sequencing was performed to validate the largest spliced sequence from 8 clusters, resulting in the confirmation of five sequences as true human transcript fragments. Some of them were differentially expressed between tumor and normal tissue by an in silico analysis. We can conclude that after clean up of the no-match dataset, we still have about 939 new exons and 165 unpredicted genes that could complete the prostate or breast transcriptome.

Breast Neoplasms↗

The UCSC Archaeal Genome Browser.

As more archaeal genomes are sequenced, effective research and analysis tools are needed to integrate the diverse information available for any given locus. The feature-rich UCSC Genome Browser, created originally to annotate the human genome, can be applied to any sequenced organism. We have created a UCSC Archaeal Genome Browser, available at http://archaea.ucsc.edu/, currently with 26 archaeal genomes. It displays G/C content, gene and operon annotation from multiple sources, sequence motifs (promoters and Shine-Dalgarno), microarray data, multi-genome alignments and protein conservation across phylogenetic and habitat categories. We encourage submission of new experimental and bioinformatic analysis from contributors. The purpose of this tool is to aid biological discovery and facilitate greater collaboration within the archaeal research community.

Archaeal Proteins↗

coliBASE: an online database for Escherichia coli, Shigella and Salmonella comparative genomics.

We have constructed coliBASE, a database for Escherichia coli, Shigella and Salmonella comparative genomics available online at http://colibase. bham.ac.uk. Unlike other E.coli databases, which focus on the laboratory model strain K12, coliBASE is intended to reflect the full diversity of E.coli and its relatives. The database contains comparative data including whole genome alignments and lists of putative orthologous genes, together with numerous analytical tools and links to existing online resources. The data are stored in a relational database, accessible by a number of user-friendly search methods and graphical browsers. The database schema is generic and can easily be applied to other bacterial genomes. Two such databases, CampyDB (for the analysis of Campylobacter spp.) and ClostriDB (for Clostridium spp.) are also available at http://campy.bham.ac.uk and http://clostri. bham.ac.uk, respectively. An example of the power of E.coli comparative analyses such as those available through coliBASE is presented.

Computational Biology↗

Genome image programs: visualization and interpretation of Escherichia coli microarray experiments.

We have developed programs to facilitate analysis of microarray data in Escherichia coli. They fall into two categories: manipulation of microarray images and identification of known biological relationships among lists of genes. A program in the first category arranges spots from glass-slide DNA microarrays according to their position in the E. coli genome and displays them compactly in genome order. The resulting genome image is presented in a web browser with an image map that allows the user to identify genes in the reordered image. Another program in the first category aligns genome images from two or more experiments. These images assist in visualizing regions of the genome with common transcriptional control. Such regions include multigene operons and clusters of operons, which are easily identified as strings of adjacent, similarly colored spots. The images are also useful for assessing the overall quality of experiments. The second category of programs includes a database and a number of tools for displaying biological information about many E. coli genes simultaneously rather than one gene at a time, which facilitates identifying relationships among them. These programs have accelerated and enhanced our interpretation of results from E. coli DNA microarray experiments. Examples are given.

Chromosome Mapping↗

ECRbase: database of evolutionary conserved regions, promoters, and transcription factor binding sites in vertebrate genomes.

Evolutionary conservation of DNA sequences provides a tool for the identification of functional elements in genomes. We have created a database of evolutionary conserved regions (ECRs) in vertebrate genomes, entitled ECRbase, which is constructed from a collection of whole-genome alignments produced by the ECR Browser. ECRbase features a database of syntenic blocks that recapitulate the evolution of rearrangements in vertebrates and a comprehensive collection of promoters in all vertebrate genomes generated using multiple sources of gene annotation. The database also contains a collection of annotated transcription factor binding sites (TFBSs) in evolutionary conserved and promoter elements. ECRbase currently includes human, rhesus macaque, dog, opossum, rat, mouse, chicken, frog, zebrafish and fugu genomes. It is freely accessible at http://ecrbase.dcode.org.

Animals↗

Diversity of the genus Lactobacillus revealed by comparative genomics of five species.

The genus Lactobacillus contains over 80 recognized species, and is characterized by a high level of diversity, reflected in its complex phylogeny. The authors' recent determination of the genome sequence of Lactobacillus salivarius means that five complete genomes of Lactobacillus species are available for comparative genomics: L. salivarius, L. plantarum, L. acidophilus, L. johnsonii and L. sakei. This paper now shows that there is no extensive synteny of the genome sequences of these five lactobacilli. Phylogeny based on whole-genome alignments suggested that L. salivarius was closer to L. plantarum than to L. sakei, which was closest to Enterococcus faecalis, in contrast to 16S rRNA gene relatedness. A total of 593 orthologues common to all five species were identified. Species relatedness based on this protein set was largely concordant with genome synteny-based relatedness. A Lactobacillus supertree, combining individual phylogenetic trees from each of 354 core proteins, had four main branches, comprising L. salivarius-L. plantarum; L. sakei; E. faecalis; and L. acidophilus-L. johnsonii. The extreme divergence of the Lactobacillus genomes analysed supports the recognition of new subgeneric divisions.

Genes, Bacterial↗

Scale-invariant structure of strongly conserved sequence in genomic intersections and alignments.

A power-law distribution of the length of perfectly conserved sequence from mouse/human whole-genome intersection and alignment is exhibited. Spatial correlations of these elements within the mouse genome are studied. It is argued that these power-law distributions and correlations are comprised in part by functional noncoding sequence and ought to be accounted for in estimating the statistical significance of apparent sequence conservation. These inter-genomic correlations of conservation are placed in the context of previously observed intra-genomic correlations, and their possible origins and consequences are discussed.

Animals↗

Mammalian small nucleolar RNAs are mobile genetic elements.

Small nucleolar RNAs (snoRNAs) of the H/ACA box and C/D box categories guide the pseudouridylation and the 2'-O-ribose methylation of ribosomal RNAs by forming short duplexes with their target. Similarly, small Cajal body-specific RNAs (scaRNAs) guide modifications of spliceosomal RNAs. The vast majority of vertebrate sno/scaRNAs are located in introns of genes transcribed by RNA polymerase II and processed by exonucleolytic trimming after splicing. A bioinformatic search for orthologues of human sno/scaRNAs in sequenced mammalian genomes reveals the presence of species- or lineage-specific sno/scaRNA retroposons (sno/scaRTs) characterized by an A-rich tail and an approximately 14-bp target site duplication that corresponds to their insertion site, as determined by interspecific genomic alignments. Three classes of snoRTs are defined based on the extent of intron and exon sequences from the snoRNA parental host gene they contain. SnoRTs frequently insert in gene introns in the sense orientation at genomic hot spots shared with other genetic mobile elements. Previously characterized human snoRNAs are encoded in retroposons whose parental copies can be identified by phylogenic analysis, showing that snoRTs can be faithfully processed. These results identify snoRNAs as a new family of mobile genetic elements. The insertion of new snoRNA copies might constitute a safeguard mechanism by which the biological activity of snoRNAs is maintained in spite of the risk of mutations in the parental copy. I furthermore propose that retroposition followed by genetic drift is a mechanism that increased snoRNA diversity during vertebrate evolution to eventually acquire new RNA-modification functions.

Animals↗

Interpreting mammalian evolution using Fugu genome comparisons.

Recently, it has been shown that a significant number of evolutionarily conserved human-Fugu noncoding elements function as tissue-specific transcriptional enhancers in vivo, suggesting that distant comparisons are capable of identifying a particular class of regulatory elements. We therefore hypothesized that by juxtaposing human/Fugu and human/mouse conservation patterns we can define conservation criteria for discovering transcriptional regulatory elements specific to mammals. Genome-scale comparisons of noncoding human/Fugu evolutionary conserved elements (ECRs) and their humans/mouse counterparts revealed a particular signature common to human/mouse ECRs (>or=350 bp long, >or=77% identity) that are also conserved in fishes. This newly defined threshold identifies 90% of all human/Fugu noncoding ECRs without the assistance of human-Fugu genome alignments and provides a very efficient filter for identifying functional human/mouse ECRs.

Animals↗

Strong and weak male mutation bias at different sites in the primate genomes: insights from the human-chimpanzee comparison.

Male mutation bias is a higher mutation rate in males than in females thought to result from the greater number of germ line cell divisions in males. If errors in DNA replication cause most mutations, then the magnitude of male mutation bias, measured as the male-to-female mutation rate ratio (alpha), should reflect the relative excess of male versus female germ line cell divisions. Evolutionary rates averaged among all sites in a sequence and compared between mammalian sex chromosomes were shown to be indeed higher in males than in females. However, it is presently unknown whether individual classes of substitutions exhibit such bias. To address this issue, we investigated male mutation bias separately at non-CpG and CpG sites using human-chimpanzee whole-genome alignments. We observed strong male mutation bias at non-CpG sites: alpha in the X-autosome comparison was approximately 6-7, which was similar to the male-to-female ratio in the number of germ line cell divisions. In contrast, mutations at CpG sites exhibited weak male mutation bias: alpha in the X-autosome comparison was only approximately 2-3. This is consistent with the methylation-induced and replication-independent mechanism of CpG transitions, which constitute the majority of mutations at CpG sites. Interestingly, our study also indicated weak male mutation bias for transversions at CpG sites, implying a spontaneous mechanism largely not associated with replication. Male mutation bias was equally strong at CpG and non-CpG sites located within unmethylated "CpG islands," suggesting the replication-dependent origin of these mutations. Thus, we found that the strength of male mutation bias is nonuniform in the primate genomes. Importantly, we discovered that male mutation bias depends on the proportion of CpG sites in the loci compared. This might explain the differences in the magnitude of primate male mutation bias observed among studies.

Animals↗

Genome-wide analysis of mammalian DNA segment fusion/fission.

As a powerful tool for gene function prediction, gene fusion has been widely studied in prokaryotes and certain groups of eukaryotes, but it has been little applied in studies of mammalian genomes. With the first fully sequenced mammalian genomes (human, mouse, rat) now available, we defined and collected a set of fusion/fission event-linked segments (FFLS) based on structured organized genomic alignment. The statistics of the sequence features highlighted the FFLSs against their random context. We found that there are three groups of FFLSs with different component pairs (i.e. gene-gene, gene-noncoding and noncoding-noncoding) in all three mammalian genomes. The proteins encoded by the components of FFLSs in the first group shown a strong tendency to interact with each other. The segmental components in the last two groups which did not contain any protein-coding genes, were found not only to be transcribed to some level, but also more conserved than the random background. Thus, these segments are possibly carrying certain biologically functional elements. We propose that FFLS may be a potential tool for prediction and analysis of function and functional interaction of genetic elements, including both genes and noncoding elements, in mammalian genomes. The full list of the FFLSs in the genomes of the three mammals is available as supporting information at doi:10.1016/j.jtbi.2005.09.016.

Animals↗

GeneOrder3.0: software for comparing the order of genes in pairs of small bacterial genomes.

BACKGROUND: An increasing number of whole viral and bacterial genomes are being sequenced and deposited in public databases. In parallel to the mounting interest in whole genomes, the number of whole genome analyses software tools is also increasing. GeneOrder was originally developed to provide an analysis of genes between two genomes, allowing visualization of gene order and synteny comparisons of any small genomes. It was originally developed for comparing virus, mitochondrion and chloroplast genomes. This is now extended to small bacterial genomes of sizes less than 2 Mb. RESULTS: GeneOrder3.0 has been developed and validated successfully on several small bacterial genomes (ca. 580 kb to 1.83 Mb) archived in the NCBI GenBank database. It is an updated web-based "on-the-fly" computational tool allowing gene order and synteny comparisons of any two small bacterial genomes. Analyses of several bacterial genomes show that a large amount of gene and genome re-arrangement occurs, as seen with earlier DNA software tools. This can be displayed at the protein level using GeneOrder3.0. Whole genome alignments of genes are presented in both a table and a dot plot. This allows the detection of evolutionary more distant relationships since protein sequences are more conserved than DNA sequences. CONCLUSIONS: GeneOrder3.0 allows researchers to perform comparative analysis of gene order and synteny in genomes of sizes up to 2 Mb "on-the-fly." AVAILABILITY: http://binf.gmu.edu/genometools.html and http://pasteur.atcc.org:8050/GeneOrder3.0.

Chromosome Mapping↗

The mechanism of cytoplasmic orthopoxvirus DNA replication.

Orthopoxvirus DNA replication occurs in the cytoplasm of infected cells within discrete foci designated as virosomes. We show that newly synthesized rabbit poxvirus (RPV) virosomal DNA consists predominantly of concatamers wherein unit length molecules are joined by fusion of two left (LL) or right (RR) ends, resulting in genomes aligned in alternating head-to-head and tail-to tail mirror image arrays. These concatameric molecules serve as the substrates from which unit length DNA molecules are excised during morphogenesis. We propose a mechanism by which internal deletions within these concatameric arrays prior to genome excision and packaging could create inverted terminal repeats and generate gene duplications.

Cytoplasm↗