Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Comparative organization of wheat homoeologous group 3S and 7L using wheat-rice synteny and identification of potential markers for genes controlling xanthophyll content in wheat.

EST and genomic DNA sequencing efforts for rice and wheat have provided the basis for interpreting genome organization and evolution. In this study we have used EST and genomic sequencing information and a bioinformatic approach in a two-step strategy to align portions of the wheat and rice genomes. In the first step, wheat ESTs were used to identify rice orthologs and it was shown that wheat 3S and rice 1 contain syntenic units with intrachromosomal rearrangements. Further analysis using anchored rice contiguous sequences and TBLASTX alignments in a second alignment step showed interruptions by orthologous genes that map elsewhere in the wheat genome. This indicates that gene content and order is not as conserved as large chromosomal blocks as previously predicted. Similarly, chromosome 7L contains syntenic units with rice 6 and 8 but is interrupted by combinations of intrachromosomal and interchromosomal rearrangements involving syntenic units and single gene orthologs from other rice chromosome groups. We have used the rice sequence annotations to identify genes that can be used to develop markers linked to biosynthetic pathways on 3BS controlling xanthophyll production in wheat and thus involved in determining flour colour.

Amino Acid Sequence↗

Significant associations of the mitochondrial transcription factor A promoter polymorphisms with marbling and subcutaneous fat depth in Wagyu x Limousin F2 crosses.

Mitochondrial transcription factor A (TFAM), a nucleus-encoded protein, regulates the initiation of transcription and replication of mitochondrial DNA (mtDNA). Decreased expression of nuclear-encoded mitochondrial genes has been associated with onset of obesity in mice. Therefore, we hypothesized genetic variants in TFAM gene influence mitochondrial biogenesis consequently affecting body fat deposition and energy metabolism. In the present study, both cDNA (2259 bp) and genomic DNA (16,666 bp) sequences were generated for the bovine TFAM gene using a combination of in silico cloning with targeted region PCR amplification. Alignment of both cDNA and genomic sequences led to the determination of genomic organization and characterization of the promoter region of the bovine TFAM gene. Two closely linked A/C and C/T single nucleotide polymorphisms (SNPs) were found in the bovine TFAM promoter and then genotyped on 237 Wagyu x Limousin F(2) animals with recorded phenotypes for marbling and subcutaneous fat depth (SFD). Statistical analysis demonstrated that both SNPs and their haplotypes were associated with marbling (P=0.0153 for A/C, P=0.0026 for C/T, and P=0.0004 for haplotype) and SFD (P=0.0200 for A/C, P=0.0039 for C/T, and P=0.0029 for haplotype), respectively. A search for transcriptional regulatory elements using MatInspector indicated that both SNPs lead to a gain/loss of six putative-binding sites for transcription factors relevant to fat deposition and energy metabolism. Our results suggest for the first time that TFAM gene plays an important role in lipid metabolism and may be a strong candidate gene for obesity in mammals.

3T3-L1 Cells↗

Inferences on the genome structure of progenitor maize through comparative analysis of rice, maize and the domesticated panicoids.

Corn and rice genetic linkage map alignments were extended and refined by the addition of 262 new, reciprocally mapped maize cDNA loci. Twenty chromosomal rearrangements were identified in maize relative to rice and these included telomeric fusions between rice linkage groups, nested insertion of rice linkage groups, intrachromosomal inversions, and a nonreciprocal translocation. Maize genome evolution was inferred relative to other species within the Panicoideae and a progenitor maize genome with eight linkage groups was proposed. Conservation of composite linkage groups indicates that the tetrasomic state arose during maize evolution either from duplication of one progenitor corn genome (autoploidy) or from a cross between species that shared the composite linkages observed in modern maize (alloploidy). New evidence of a quadruplicated homeologous segment on maize chromosomes 2 and 10, and 3 and 4, corresponded to the internally duplicated region on rice chromosomes 11 and 12 and suggested that this duplication in the rice genome predated the divergence of the Panicoideae and Oryzoideae subfamilies. Charting of the macroevolutionary steps leading to the modern maize genome clarifies the interpretation of intercladal comparative maps and facilitates alignments and genomic cross-referencing of genes and phenotypes among grass family members.

Chromosome Aberrations↗

Identification of the binding sites of regulatory proteins in bacterial genomes.

We present an algorithm that extracts the binding sites (represented by position-specific weight matrices) for many different transcription factors from the regulatory regions of a genome, without the need for delineating groups of coregulated genes. The algorithm uses the fact that many DNA-binding proteins in bacteria bind to a bipartite motif with two short segments more conserved than the intervening region. It identifies all statistically significant patterns of the form W(1)N(x)W(2), where W(1) and W(2) are two short oligonucleotides separated by x arbitrary bases, and groups them into clusters of similar patterns. These clusters are then used to derive quantitative recognition profiles of putative regulatory proteins. For a given cluster, the algorithm finds the matching sequences plus the flanking regions in the genome and performs a multiple sequence alignment to derive position-specific weight matrices. We have analyzed the Escherichia coli genome with this algorithm and found approximately 1,500 significant patterns, which give rise to approximately 160 distinct position-specific weight matrices. A fraction of these matrices match the binding sites of one-third of the approximately 60 characterized transcription factors with high statistical significance. Many of the remaining matrices are likely to describe binding sites and regulons of uncharacterized transcription factors. The significance of these matrices was evaluated by their specificity, the location of the predicted sites, and the biological functions of the corresponding regulons, allowing us to suggest putative regulatory functions. The algorithm is efficient for analyzing newly sequenced bacterial genomes for which little is known about transcriptional regulation.

Algorithms↗

Pigs in sequence space: a 0.66X coverage pig genome survey based on shotgun sequencing.

BACKGROUND: Comparative whole genome analysis of Mammalia can benefit from the addition of more species. The pig is an obvious choice due to its economic and medical importance as well as its evolutionary position in the artiodactyls. RESULTS: We have generated approximately 3.84 million shotgun sequences (0.66X coverage) from the pig genome. The data are hereby released (NCBI Trace repository with center name "SDJVP", and project name "Sino-Danish Pig Genome Project") together with an initial evolutionary analysis. The non-repetitive fraction of the sequences was aligned to the UCSC human-mouse alignment and the resulting three-species alignments were annotated using the human genome annotation. Ultra-conserved elements and miRNAs were identified. The results show that for each of these types of orthologous data, pig is much closer to human than mouse is. Purifying selection has been more efficient in pig compared to human, but not as efficient as in mouse, and pig seems to have an isochore structure most similar to the structure in human. CONCLUSION: The addition of the pig to the set of species sequenced at low coverage adds to the understanding of selective pressures that have acted on the human genome by bisecting the evolutionary branch between human and mouse with the mouse branch being approximately 3 times as long as the human branch. Additionally, the joint alignment of the shot-gun sequences to the human-mouse alignment offers the investigator a rapid way to defining specific regions for analysis and resequencing.

Animals↗

Alignment of metabolic pathways.

MOTIVATION: Several genome-scale efforts are underway to reconstruct metabolic networks for a variety of organisms. As the resulting data accumulates, the need for analysis tools increases. A notable requirement is a pathway alignment finder that enables both the detection of conserved metabolic pathways among different species as well as divergent metabolic pathways within a species. When comparing two pathways, the tool should be powerful enough to take into account both the pathway topology as well as the nodes' labels (e.g. the enzymes they denote), and allow flexibility by matching similar--rather than identical--pathways. RESULTS: MetaPathwayHunter is a pathway alignment tool that, given a query pathway and a collection of pathways, finds and reports all approximate occurrences of the query in the collection, ranked by similarity and statistical significance. It is based on a novel, efficient graph matching algorithm that extends the functionality of known techniques. The program also supports a visualization interface with which the alignment of two homologous pathways can be graphically displayed. We employed this tool to study the similarities and differences in the metabolic networks of the bacterium Escherichia coli and the yeast Saccharomyces cerevisiae, as represented in highly curated databases. We reaffirmed that most known metabolic pathways common to both the species are conserved. Furthermore, we discovered a few intriguing relationships between pathways that provide insight into the evolution of metabolic pathways. We conclude with a description of biologically meaningful meta-queries, demonstrating the power and flexibility of our new tool in the analysis of metabolic pathways.

Animals↗

Complete nucleotide sequence of genotype 4 hepatitis C viruses isolated from patients co-infected with human immunodeficiency virus type 1.

The hepatitis C virus (HCV) genotype 4 is spreading among southern European intravenous drug users, who are frequently co-infected with human immunodeficiency virus type 1 (HIV-1). Response to interferon (IFN) alpha-based therapies in HIV-1 positive patients co-infected with HCV genotype 4 is poor, similar to that obtained for HCV genotype 1 and much lower than for HCV genotypes 2 and 3. The lack of sequence data related to HCV of genotype 4 prompted us to sequence the complete genome of two genotype 4 variants isolated from two HIV-1 co-infected patients (24 and 25). Our aim was to investigate the evolutionary relationships of the former variants with other genotypes and/or genotype 4 subtypes. Sequence alignments and phylogenetic analysis from genomic regions 5'NC, core-E1 and NS5B revealed that the variants isolated from patients 24 and 25 (both subtyped 4c/4d by INNO-LIPA II HCV) belong to subtypes 4d and 4a, respectively. When looking at the complete genome sequence one of the variants showed a new genotype 4 subtype. Interestingly, sequence length differences in the interferon sensitivity determining region coding regions were observed when compared with sequences from other genotypes. Similarly, when the catalytic efficiency of the NS3/4 protease from patients 24 and 25 samples were determined, they displayed 70.6+/-7.7 and 23.5+/-3.4%, respectively, of the activity shown by genotype 1 NS3/4 proteases. Overall, pairwise comparison and phylogenetic analysis of nucleotide sequences of the complete genome or the different protein-encoding regions showed that genotype 4 sequences were more closely related to genotype 1 sequences. The description of new HCV genome variants may help our understanding of the HCV biology as well as the role of different genotypes in HCV treatment and therapy response.

Amino Acid Sequence↗

Retrieval and on-the-fly alignment of sequence fragments from the HIV database.

MOTIVATION: The amount of HIV-1 sequence data generated (presently around 42000 sequences, of which more than 22000 are from the V3 region of the viral envelope) presents a challenge for anyone working on the analysis of these data. A major problem is obtaining the region of interest from the stored sequences, which often contain but are not limited to that region. In addition, multiple alignment programs generally cannot deal with the large numbers of sequences that are available for many HIV-1 regions. We set out to provide our users with a tool that will retrieve and create an initial alignment of the HIV sequences that are available for a given genomic region. RESULTS: The MPAlign (Multiple Pairwise Alignment) web interface is a collection of Perl scripts that retrieves sequences from the Los Alamos HIV sequence database based on a number of search parameters. All sequences were pairwise-aligned to a model sequence using the Hidden Markov Model-based program HMMER. The HMMER model is general enough to accommodate virtually all HIV-1 sequences stored in the database. To create a multiple sequence alignment, gaps were inserted into the sequences during retrieval, so that they are aligned to one another. Retrieving and aligning the almost 560 gp120 sequences (approximately>1500 nt) stored in the database is at least 1500 times faster than a similar Clustal alignment.

Algorithms↗

Phytophthora functional genomics database (PFGD): functional genomics of phytophthora-plant interactions.

The Phytophthora Functional Genomics Database (PFGD; http://www.pfgd.org), developed by the National Center for Genome Resources in collaboration with The Ohio State University-Ohio Agricultural Research and Development Center (OSU-OARDC), is a publicly accessible information resource for Phytophthora-plant interaction research. PFGD contains transcript, genomic, gene expression and functional assay data for Phytophthora infestans, which causes late blight of potato, and Phytophthora sojae, which affects soybeans. Automated analyses are performed on all sequence data, including consensus sequences derived from clustered and assembled expressed sequence tags. The PFGD search filter interface allows intuitive navigation of transcript and genomic data organized by library and derived queries using modifiers, annotation keywords or sequence names. BLAST services are provided for libraries built from the transcript and genomic sequences. Transcript data visualization tools include Quality Screening, Multiple Sequence Alignment and Features and Annotations viewers. A genomic browser that supports comparative analysis via novel dynamic functional annotation comparisons is also provided. PFGD is integrated with the Solanaceae Genomics Database (SolGD; http://www.solgd.org) to help provide insight into the mechanisms of infection and resistance, specifically as they relate to the genus Phytophthora pathogens and their plant hosts.

Algal Proteins↗

Assessing the level of collinearity between Arabidopsis thaliana and Brassica napus for A. thaliana chromosome 5.

This study describes a comprehensive comparison of chromosome 5 of the model crucifer Arabidopsis with the genome of its amphidiploid crop relative Brassica napus and introduces the use of in silico sequence homology to identify conserved loci between the two species. A region of chromosome 5, spanning 8 Mb, was found in six highly conserved copies in the B. napus genome. A single inversion appeared to be the predominant rearrangement that had separated the two lineages leading to the formation of Arabidopsis chromosome 5 and its homologues in B. napus. The observed results could be explained by the fusion of three ancestral genomes with strong similarities to modern-day Arabidopsis to generate the constituent diploid genomes of B. napus. This supports the hypothesis that the diploid Brassica genomes evolved from a common hexaploid ancestor. Alignment of the genetic linkage map of B. napus with the genomic sequence of Arabidopsis indicated that for specific regions a genetic distance of 1 cM in B. napus was equivalent to 285 Kb of Arabidopsis DNA sequence. This analysis strongly supports the application of Arabidopsis as a tool in marker development, map-based gene cloning, and candidate gene identification for the larger genomes of Brassica crop species.

Arabidopsis↗

Large-scale Homologous Analysis of Genome Sequence.

We described a new method for the large-scale homologous analysis of genome sequences, which used hashing technique combined with sparse dynamic programming to get a sequence alignment. Three examples, the plant chloroplast genomes, the mammalian T-cell receptors C(alpha)/C(delta) gene loci and the mammalian gamma-crystallin gene clusters were analysed. The results showed that the method was more rapid to obtain accurate enough data and might be useful in genome analysis.

Journal Article↗

Current bioinformatics tools in genomic biomedical research (Review).

On the advent of a completely assembled human genome, modern biology and molecular medicine stepped into an era of increasingly rich sequence database information and high-throughput genomic analysis. However, as sequence entries in the major genomic databases currently rise exponentially, the gap between available, deposited sequence data and analysis by means of conventional molecular biology is rapidly widening, making new approaches of high-throughput genomic analysis necessary. At present, the only effective way to keep abreast of the dramatic increase in sequence and related information is to apply biocomputational approaches. Thus, over recent years, the field of bioinformatics has rapidly developed into an essential aid for genomic data analysis and powerful bioinformatics tools have been developed, many of them publicly available through the World Wide Web. In this review, we summarize and describe the basic bioinformatics tools for genomic research such as: genomic databases, genome browsers, tools for sequence alignment, single nucleotide polymorphism (SNP) databases, tools for ab initio gene prediction, expression databases, and algorithms for promoter prediction.

Computational Biology↗

Data integration and genomic medicine.

Genomic medicine aims to revolutionize health care by applying our growing understanding of the molecular basis of disease. Research in this arena is data intensive, which means data sets are large and highly heterogeneous. To create knowledge from data, researchers must integrate these large and diverse data sets. This presents daunting informatic challenges such as representation of data that is suitable for computational inference (knowledge representation), and linking heterogeneous data sets (data integration). Fortunately, many of these challenges can be classified as data integration problems, and technologies exist in the area of data integration that may be applied to these challenges. In this paper, we discuss the opportunities of genomic medicine as well as identify the informatics challenges in this domain. We also review concepts and methodologies in the field of data integration. These data integration concepts and methodologies are then aligned with informatics challenges in genomic medicine and presented as potential solutions. We conclude this paper with challenges still not addressed in genomic medicine and gaps that remain in data integration research to facilitate genomic medicine.

Biomedical Research↗

ASAP: the Alternative Splicing Annotation Project.

Recently, genomics analyses have demonstrated that alternative splicing is widespread in mammalian genomes (30-60% of genes reported to have multiple isoforms), and may be one of their most important mechanisms of functional regulation. However, by comparison with other genomics data such as genome annotation, SNPs, or gene expression, there exists relatively little database infrastructure for the study of alternative splicing. We have constructed an online database ASAP (the Alternative Splicing Annotation Project) for biologists to access and mine the enormous wealth of alternative splicing information coming from genomics and proteomics. ASAP is based on genome-wide analyses of alternative splicing in human (30 793 alternative splice relationships found) from detailed alignment of expressed sequences onto the genomic sequence. ASAP provides precise gene exon-intron structure, alternative splicing, tissue specificity of alternative splice forms, and protein isoform sequences resulting from alternative splicing. Moreover, it can help biologists design probe sequences for distinguishing specific mRNA isoforms. ASAP is intended to be a community resource for collaborative annotation of alternative splice forms, their regulation, and biological functions. The URL for ASAP is http://www.bioinformatics.ucla.edu/ASAP.

Alternative Splicing↗

Peptidylprolyl cis/trans isomerases (immunophilins): biological diversity--targets--functions.

Information recovered from genome sequencing projects, multiple sequence alignments, structural analyses of PPIase and published records were used in deciphering the biological diversity, functions and targets of four groups of proteins encoded by dissimilar sets of sequences whose spatial representations exhibit peptidylprolyl cis/trans isomerase activity (PPIase). In the human genome there are encoded fifteen proteins whose segments have significant homology with the sequence of 12 kDa protein which is the target of the potent immunosuppressive macrolides FK506 or rapamycin. The 12 kDa archetype of the FK506-binding protein (FKBP), known as FKBP-12a, is an abundant intracellular protein whereas other FKBPs possessing from one to four FK506-like binding domains (FKBDs) have nominal masses varying from 13 to 135 kDa. The human genome contains at least sixteen genes encoding proteins comprising one cyclosporin-A (CsA) binding domain (CLD) called cyclophilins whose nominal masses vary from 17 to 324 kDa and multiple coding segments for small cyclophilins (17-19 kDa) whose transcription levels and functions remain unknown. The third group of PPIases encoded in the genome comprises two proteins (hPin1 and hParv14) where hPin1 is an important PPIase for cell cycle. The A. thaliana, C. elegans, D. melanogaster and S. cerevisiae genomes encode a less diverse spectrum of PPIases whereas the prokaryotic genomes contain from none to three cyclophilins, from none to four genes encoding FKBPs, one distant homologue of the Pin1 protein named parvulin and the fourth group of PPIases known as trigger factors. PPIases are discretely distributed to different cellular compartments and interact with a number of targets that control a range of cellular processes. Analyses of the sequence alignments of the two groups of PPIases, namely cyclophilins and FKBPs from diverse phyla, show that in each group their sequences diverge but the amino acid residues which form the PPIase activity site and macrolide binding cavity remain well conserved in the majority of them which suggests that the spatial structures and functions of each group of PPIases remain conserved.

Amino Acid Sequence↗

MAVG: locating non-overlapping maximum average segments in a given sequence.

SUMMARY: MAVG is a software tool for finding k non-overlapping maximum-average segments that are sufficiently long in a given sequence of real numbers, for any k > 0. It has applications in several areas of biomolecular sequence analysis including locating GC-rich regions and CpG islands in a genomic sequence, and annotating multiple sequence alignments. AVAILABILITY: http://iubio.bio.indiana.edu/soft/molbio/pattern/cpg_islands/.

Algorithms↗

Genome BLAST distance phylogenies inferred from whole plastid and whole mitochondrion genome sequences.

BACKGROUND: Phylogenetic methods which do not rely on multiple sequence alignments are important tools in inferring trees directly from completely sequenced genomes. Here, we extend the recently described Genome BLAST Distance Phylogeny (GBDP) strategy to compute phylogenetic trees from all completely sequenced plastid genomes currently available and from a selection of mitochondrial genomes representing the major eukaryotic lineages. BLASTN, TBLASTX, or combinations of both are used to locate high-scoring segment pairs (HSPs) between two sequences from which pairwise similarities and distances are computed in different ways resulting in a total of 96 GBDP variants. The suitability of these distance formulae for phylogeny reconstruction is directly estimated by computing a recently described measure of "treelikeness", the so-called delta value, from the respective distance matrices. Additionally, we compare the trees inferred from these matrices using UPGMA, NJ, BIONJ, FastME, or STC, respectively, with the NCBI taxonomy tree of the taxa under study. RESULTS: Our results indicate that, at this taxonomic level, plastid genomes are much more valuable for inferring phylogenies than are mitochondrial genomes, and that distances based on breakpoints are of little use. Distances based on the proportion of "matched" HSP length to average genome length were best for tree estimation. Additionally we found that using TBLASTX instead of BLASTN and, particularly, combining TBLASTX and BLASTN leads to a small but significant increase in accuracy. Other factors do not significantly affect the phylogenetic outcome. The BIONJ algorithm results in phylogenies most in accordance with the current NCBI taxonomy, with NJ and FastME performing insignificantly worse, and STC performing as well if applied to high quality distance matrices. delta values are found to be a reliable predictor of phylogenetic accuracy. CONCLUSION: Using the most treelike distance matrices, as judged by their delta values, distance methods are able to recover all major plant lineages, and are more in accordance with Apicomplexa organelles being derived from "green" plastids than from plastids of the "red" type. GBDP-like methods can be used to reliably infer phylogenies from different kinds of genomic data. A framework is established to further develop and improve such methods. delta values are a topology-independent tool of general use for the development and assessment of distance methods for phylogenetic inference.

Algorithms↗