Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 955 records · Page 53Linked to original sources

Phylogenetic and structural analyses of the oxa1 family of protein translocases.

Mitochondrial Oxa1p homologs have been shown to function in protein export and membrane insertion in bacteria, mitochondria and chloroplasts, but their mode of action, organismal distribution and evolutionary origins are poorly understood. All sequenced homologs of Oxa1p were retrieved from the databases and multiply aligned. All organisms with a fully sequenced genome possess at least one Oxa1p homolog showing that the family is truly ubiquitous. Most prokaryotes possess just one Oxa1p homolog, but several Gram-positive bacteria and one archaeon possess two, and eukaryotes may have as many as six. Although these proteins vary in length over a 5-fold range, they exhibit a common hydrophobic core region of about 200 residues. Multiple sequence alignments reveal conserved residues and provide the basis for structural and phylogenetic analyses that serve to characterize the Oxa1 family.

Amino Acid Sequence↗

MACSIMS: multiple alignment of complete sequences information management system.

BACKGROUND: In the post-genomic era, systems-level studies are being performed that seek to explain complex biological systems by integrating diverse resources from fields such as genomics, proteomics or transcriptomics. New information management systems are now needed for the collection, validation and analysis of the vast amount of heterogeneous data available. Multiple alignments of complete sequences provide an ideal environment for the integration of this information in the context of the protein family. RESULTS: MACSIMS is a multiple alignment-based information management program that combines the advantages of both knowledge-based and ab initio sequence analysis methods. Structural and functional information is retrieved automatically from the public databases. In the multiple alignment, homologous regions are identified and the retrieved data is evaluated and propagated from known to unknown sequences with these reliable regions. In a large-scale evaluation, the specificity of the propagated sequence features is estimated to be >99%, i.e. very few false positive predictions are made. MACSIMS is then used to characterise mutations in a test set of 100 proteins that are known to be involved in human genetic diseases. The number of sequence features associated with these proteins was increased by 60%, compared to the features available in the public databases. An XML format output file allows automatic parsing of the MACSIM results, while a graphical display using the JalView program allows manual analysis. CONCLUSION: MACSIMS is a new information management system that incorporates detailed analyses of protein families at the structural, functional and evolutionary levels. MACSIMS thus provides a unique environment that facilitates knowledge extraction and the presentation of the most pertinent information to the biologist. A web server and the source code are available at http://bips.u-strasbg.fr/MACSIMS/.

Algorithms↗

Pattern of selective constraint in C. elegans and C. briggsae genomes.

Similarity between related genomes may carry information on selective constraint in each of them. We analysed patterns of similarity between several homologous regions of Caenorhabditis elegans and C. briggsae genomes. All homologous exons are quite similar. Alignments of introns and of intergenic sequences contain long gaps, segments where similarity is low and close to that between random sequences aligned using the same parameters, and segments of high similarity. Conservative estimates of the fractions of selectively constrained nucleotides are 72%, 17% and 18% for exons, introns and intergenic sequences, respectively. This implies that the total number of constrained nucleotides within non-coding sequences is comparable to that within coding sequences, so that at least one-third of nucleotides in C. elegans and C. briggsae genomes are under strong stabilizing selection.

Animals↗

HIV may produce inhibitory microRNAs (miRNAs) that block production of CD28, CD4 and some interleukins.

It is well-known that HIV-1 infection results in a gradual decline of the CD4+ T-lymphocytes, but the underlying mechanism of this decline is not completely understood. Research has shown that HIV-1 infection of CD4+ T cells results in decreased CD28 expression, but the mechanism of this repression is unknown. There is also substantial evidence demonstrating regulatory involvement of microRNA (miRNA) during protein expression in plants and some animals, and reports have recently been published confirming the existence of viral-encoded miRNAs. Based on these findings, we hypothesize that viral-encoded miRNA from HIV-1 may directly alter T cell, macrophage and dendritic cell activity. To investigate a potential correlation between the genomic complementarity of HIV-1 and host cell protein expression, a local alignment search was performed to assess for regions of complementarity between the HIV-1 proviral genome and the mRNA coding sequence of various proteins expressed by CD+ T cells and macrophages. Regions of complementarity with strong correlations to the currently established criteria for miRNA:target mRNA activity were found between HIV-1 and CD28, CTLA-4 and some interleukins, suggesting that HIV-1 may produce translational repression in host cells.

Base Sequence↗

xGDB: open-source computational infrastructure for the integrated evaluation and analysis of genome features.

The eXtensible Genome Data Broker (xGDB) provides a software infrastructure consisting of integrated tools for the storage, display, and analysis of genome features in their genomic context. Common features include gene structure annotations, spliced alignments, mapping of repetitive sequence, and microarray probes, but the software supports inclusion of any property that can be associated with a genomic location. The xGDB distribution and user support utilities are available online at the xGDB project website, http://xgdb.sourceforge.net/.

Arabidopsis↗

Genome-wide subgenome-resolved analysis validates chromosome 4 differentiation and prioritizes introgressed Coffea arabica accessions.

Chromosome 4 introgression in Timor hybrid-derived Coffea arabica is established, but the robustness of accession prioritization and the relative strength of cultivated-introgressed differentiation across the canephora-derived (sgC) and eugenioides-derived (sgE) subgenomes remained unclear under explicit subgenome filtering. We reanalyzed public genomic resources from 44 coffee accessions using strict contig-level subgenome filtering, Arabica-only population-structure analysis, SNP-panel sensitivity testing, genome-wide differentiation scans, permutation testing, and direct sequence alignment. Population structure and accession rankings were stable across marker densities and random seeds, and the same six introgressed references were retained throughout. Chromosome 4 ranked first in both subgenomes, with a strong sgC signal and a markedly weaker sgE signal; independent baseline-panel permutation tests supported both chromosome 4-associated signals. Direct alignment supported correspondence to the expected chromosome 4 pseudomolecules while showing incomplete source coverage and unresolved exact boundaries. Alignment-supported blocks contained 88 sgC and 62 sgE provisional defense-, signaling-, and regulatory-associated annotations. These results provide a genome-wide, quantitatively validated framework for prioritizing introgressed germplasm and candidate chromosome 4 regions for phenotype-linked coffee research without implying equivalent introgression, exact liftover, or causal resistance genes.

Coffea arabica↗

Identification of novel highly expressed genes in pancreatic ductal adenocarcinomas through a bioinformatics analysis of expressed sequence tags.

In most microarray experiments, a significant fraction of the differentially expressed mRNAs identified correspond to expressed sequence tags (ESTs) and are generally discarded from further analyses. We used careful bioinformatics analyses to characterize those ESTs that were found to be highly overexpressed in a series of pancreatic adenocarcinomas. cDNA was prepared from 60 non-neoplastic samples (normal pancreas [n = 20], normal colon [n = 10], or normal duodenal mucosal [n = 30]) and from 64 pancreatic cancers (resected cancers [n = 50] or cancer cell lines [n = 14]) and hybridized to the complete Affymetrix Human Genome U133 GeneChip(R) set (arrays U133A and B) for simultaneous analysis of 45,000 fragments corresponding to 33,000 known genes and 6,000 ESTs. The GeneExpress(R) software system Fold Change Analysis Tool was used and 60 ESTs were identified that were expressed at levels at least 3-fold greater in the pancreatic cancers as compared to normal tissues. Searches against the human genomic sequence and comparative genomic analysis of human and mouse genomes was carried out using basic local alignment search tools (BLAST), BLASTN, and BLASTX, for identifying protein coding genes corresponding to the ESTs. Subsequently, in order to pick the most relevant candidate genes for a more detailed analysis, we looked for domains/motifs in the open reading frames using SMART and Pfam programs. We were able to definitively map 43 of the 60 ESTs to known or novel genes, and 15 of the ESTs could be localized in close proximity to a gene in the human genome although we were unable to establish that the EST was indeed derived from those genes. The differential expression of a subset of genes was confirmed at the protein level by immunohistochemical labeling of tissue microarrays (inhibin beta A [INHBA] and CD29) and/or at the transcript level by RT-PCR (INHBA, AKAP12, ELK3, FOXQ1, EIF5A2, and EFNA5). We conclude that bioinformatics tools can be used to characterize differentially overexpressed ESTs, and that some of these ESTs may represent diagnostically and therapeutically useful targets that might be missed using data solely from currently annotated databases.

Adenocarcinoma↗

MOSAIC: segmenting multiple aligned DNA sequences.

UNLABELLED: MOSAIC is a set of tools for the segmentation of multiple aligned DNA sequences into homogeneous zones. The segmentation is based on the distribution of mutational events along the alignment. As an example, the analysis of one repeated sequence belonging to the subtelomeric regions of the yeast genome is presented. AVAILABILITY: Free access from ftp://ftp.biomath.jussieu.fr/pub/papers/MOSAIC

Genome, Fungal↗

Identification of potential regulatory motifs in odorant receptor genes by analysis of promoter sequences.

Mouse odorant receptors (ORs) are encoded by >1000 genes dispersed throughout the genome. Each olfactory neuron expresses one single OR gene, while the rest of the genes remain silent. The mechanisms underlying OR gene expression are poorly understood. Here, we investigated if OR genes share common cis-regulatory sequences in their promoter regions. We carried out a comprehensive analysis in which the upstream regions of a large number of OR genes were compared. First, using RLM-RACE, we generated cDNAs containing the complete 5'-untranslated regions (5'-UTRs) for a total number of 198 mouse OR genes. Then, we aligned these cDNA sequences to the mouse genome so that the 5' structure and transcription start sites (TSSs) of the OR genes could be precisely determined. Sequences upstream of the TSSs were retrieved and browsed for common elements. We found DNA sequence motifs that are overrepresented in the promoter regions of the OR genes. Most motifs resemble O/E-like sites and are preferentially localized within 200 bp upstream of the TSSs. Finally, we show that these motifs specifically interact with proteins extracted from nuclei prepared from the olfactory epithelium, but not from brain or liver. Our results show that the OR genes share common promoter elements. The present strategy should provide information on the role played by cis-regulatory sequences in OR gene regulation.

5' Untranslated Regions↗

Structure prediction in a post-genomic environment: a secondary and tertiary structural model for the initiation factor 5A family.

Two predictions have been prepared for the fold of initiation factor 5A (IF5A) starting from a set of homologous sequences. In the first, a secondary structural model was predicted for the protein in 1994, when only eleven homologs (and no eubacterial homologs) had been sequenced. The second was made recently, after genome projects had generated a total of 33 sequences for the protein family from species of all three kingdoms of life. With the second set of sequences, but not with the first, it was possible to predict that the N-terminal domain of the protein folds in a possibly open beta-barrel/sandwich core structure, with a short helix capping one side of the barrel. We place the pair of predictions in the public domain before an experimental structure is known. This example illustrates the impact of genome sequencing projects on structure prediction from sequence alignments.

Amino Acid Sequence↗

Comparative genomics of microbial pathogens and symbionts.

We are interested in quantifying the contribution of gene acquisition, loss, expansion and rearrangements to the evolution of microbial genomes. Here, we discuss factors influencing microbial genome divergence based on pair-wise genome comparisons of closely related strains and species with different lifestyles. A particular focus is on intracellular pathogens and symbionts of the genera Rickettsia, Bartonella and BUCHNERA: Extensive gene loss and restricted access to phage and plasmid pools may provide an explanation for why single host pathogens are normally less successful than multihost pathogens. We note that species-specific genes tend to be shorter than orthologous genes, suggesting that a fraction of these may represent fossil-orfs, as also supported by multiple sequence alignments among species. The results of our genome comparisons are placed in the context of phylogenomic analyses of alpha and gamma proteobacteria. We highlight artefacts caused by different rates and patterns of mutations, suggesting that atypical phylogenetic placements can not a priori be taken as evidence for horizontal gene transfer events. The flexibility in genome structure among free-living microbes contrasts with the extreme stability observed for the small genomes of aphid endosymbionts, in which no rearrangements or inflow of genetic material have occurred during the past 50 millions years (1). Taken together, the results suggest that genomic stability correlate with the content of repeated sequences and mobile genetic elements, and thereby indirectly with bacterial lifestyles.

Alphaproteobacteria↗

A novel gene, RSD-3/HSD-3.1, encodes a meiotic-related protein expressed in rat and human testis.

The expression of stage-specific genes during spermatogenesis was determined by isolating two segments of rat seminiferous tubule at different stages of the germinal epithelium cycle delineated by transillumination-delineated microdissection, combined with differential display polymerase chain reaction to identify the differential transcripts formed. A total of 22 cDNAs were identified and accepted by GenBank as new expressed sequence tags. One of the expressed sequence tags was radiolabeled and used as a probe to screen a rat testis cDNA library. A novel full-length cDNA composed of 2228 bp, designated as RSD-3 (rat sperm DNA no.3, GenBank accession no. AF094609) was isolated and characterized. The reading frame encodes a polypeptide consisting of 526 amino acid residues, containing a number of DNA binding motifs and phosphorylation sites for PKC, CK-II, and p34cdc2. Northern blot of mRNA prepared from various tissues of adult rats showed that RSD-3 is expressed only in the testis. The initial expression of the RSD-3 gene was detected in the testis on the 30th postnatal day and attained adult level on the 60th postnatal day. Immunolocalization of RSD-3 in germ cells of rat testis showed that its expression is restricted to primary spermatocytes, undergoing meiosis division I. A human testis homologue of RSD-3 cDNA, designated as HSD-3.1 (GenBank accession no. AF144487) was isolated by screening the Human Testis Rapid-Screen arrayed cDNA library panels by RT-PCR. The exon-intron boundaries of HSD-3.1 gene were determined by aligning the cDNA sequence with the corresponding genome sequence. The cDNA consisted of 12 exons that span approximately 52.8 kb of the genome sequence and was mapped to chromosome 14q31.3.

Adult↗

Identification and complete sequencing of novel human transcripts through the use of mouse orthologs and testis cDNA sequences.

The correct identification of all human genes, and their derived transcripts, has not yet been achieved, and it remains one of the major aims of the worldwide genomics community. Computational programs suggest the existence of 30,000 to 40,000 human genes. However, definitive gene identification can only be achieved by experimental approaches. We used two distinct methodologies, one based on the alignment of mouse orthologous sequences to the human genome, and another based on the construction of a high-quality human testis cDNA library, in an attempt to identify new human transcripts within the human genome sequence. We generated 47 complete human transcript sequences, comprising 27 unannotated and 20 annotated sequences. Eight of these transcripts are variants of previously known genes. These transcripts were characterized according to size, number of exons, and chromosomal localization, and a search for protein domains was undertaken based on their putative open reading frames. In silico expression analysis suggests that some of these transcripts are expressed at low levels and in a restricted set of tissues.

Amino Acid Sequence↗

A 100-kb physical and transcriptional map around the EDH17B2 gene: identification of three novel genes and a pseudogene of a human homologue of the rat PRL-1 tyrosine phosphatase.

In this paper, we describe the physical map and transcriptional organisation of a 100-kb region with the BRCA1 locus at 17q12-21. Using the cDNA of the EDH17B2 gene as a probe, we screened a human genomic cosmid library. Positive cosmid clones were aligned and a contig around the EDH17B2 gene was established, expanding the previously reported map. In order to identify genes located in this region, we used the cosmid inserts to select cDNAs from a human ovarian cDNA library. Among the clones identified, cDNA OV-1 corresponds to a human homologue of a rat PRL-1 tyrosine phosphatase gene that shows enhanced expression during hepatic regeneration and in some tumour cell lines. Neither the OV-1 nor the PRL-1 protein shares strong homology with any previously characterised phosphotyrosine phosphatase, suggesting that they probably belong to a new phosphatase family. In an attempt to characterise the OV-1 gene, we found that the genomic sequence present on chromosome 17 probably corresponds to a nonfunctional copy of the gene, as it contains several sequence changes that disrupt the potential coding information of the gene. Three other cDNAs, corresponding to unrelated genes, were also identified and characterised. They did not reveal striking homologies in database sequence comparison and therefore represent new genes localised on chromosome 17q, in a region that frequently shows loss of heterozigosity in sporadic breast and ovarian cancers.

Amino Acid Sequence↗

Fast and reliable prediction of noncoding RNAs.

We report an efficient method for detecting functional RNAs. The approach, which combines comparative sequence analysis and structure prediction, already has yielded excellent results for a small number of aligned sequences and is suitable for large-scale genomic screens. It consists of two basic components: (i) a measure for RNA secondary structure conservation based on computing a consensus secondary structure, and (ii) a measure for thermodynamic stability, which, in the spirit of a z score, is normalized with respect to both sequence length and base composition but can be calculated without sampling from shuffled sequences. Functional RNA secondary structures can be identified in multiple sequence alignments with high sensitivity and high specificity. We demonstrate that this approach is not only much more accurate than previous methods but also significantly faster. The method is implemented in the program rnaz, which can be downloaded from www.tbi.univie.ac.at/~wash/RNAz. We screened all alignments of length n > or = 50 in the Comparative Regulatory Genomics database, which compiles conserved noncoding elements in upstream regions of orthologous genes from human, mouse, rat, Fugu, and zebrafish. We recovered all of the known noncoding RNAs and cis-acting elements with high significance and found compelling evidence for many other conserved RNA secondary structures not described so far to our knowledge.

Algorithms↗

Whole genome sequencing of unusual Hepatitis C virus subtypes and drug resistance analysis during direct-acting antiviral therapy in India.

INTRODUCTION AND OBJECTIVES: Pangenotypic direct-acting antivirals (DAA) are effective against highly prevalent Hepatitis C virus (HCV) subtypes, but have been clinically validated almost exclusively in high-income countries. Unusual HCV subtypes may carry natural polymorphisms, potentially impacting DAA susceptibility. We conducted full-genome characterization and resistance analysis of unusual HCV subtypes in patients receiving DAA treatment. PATIENTS AND METHODS: In this prospective hospital-based study, eligible patients were screened for anti-HCV antibodies and active infection was confirmed by diagnostic 5'NCR-based HCV RNA detection. Genotyping was performed by core region sequencing, and viral load quantified by real-time PCR. For whole genome sequencing, multiplex primers were designed using alignments of global reference sequences. Sequencing was carried out using the Oxford Nanopore Technology platform. Phylogenetic analysis used multiple sequence alignment and the HCV-GLUE resource for resistance-associated substitution (RAS) analysis. RESULTS: Predominant genotype was genotype 3 in 64.3% (n = 45); genotype 6 in 21.4% (n = 15); and genotype 1 in 14.2% (n = 10). Unusual HCV subtype 6xa was detected in two patients and showed no NS5A resistance mutations. One genotype 3b patient relapsed at 24 weeks post-DAA treatment completion and carried NS5A resistance-associated substitutions 30 K and 31 M both at baseline and at relapse, conferring high-level resistance to NS5A inhibitors. CONCLUSION: This is the first report from India of whole genome sequencing of HCV subtype 6xa. The identification of NS5A resistance mutations in the 3b relapse case underscores challenges for global HCV elimination strategies.

Humans↗

A novel algorithm for computational identification of contaminated EST libraries.

A key goal of the Human Genome Project was to understand the complete set of human proteins, the proteome. Since the genome sequence by itself is not sufficient for predicting new genes and alternative splicing events that lead to new proteins, expressed sequence tags (ESTs) are used as the primary tool for these purposes. The high prevalence of artifacts in dbEST, however, often leads to invalid predictions. Here we describe a novel method for recognizing genomic DNA contamination and other artifacts that cannot be identified using current EST cleaning techniques. Our method uses the alignment of the entire set of ESTs to the human genome to identify highly contaminated EST libraries. We discovered 53 highly contaminated libraries and a subset of 24 766 ESTs from these libraries that probably represent contamination with genomic DNA, pre-mRNA, and ESTs that span non-canonical introns. Although this is only a small fraction of the entire EST dataset, each contaminating sequence could create a spurious transcript prediction. Indeed, in the clustering and assembly tool that we used, these sequences would have caused incorrect inference of 9575 new splice variants and 6370 new genes. Conclusions based on EST analysis, including prediction of alternative splicing, should be re-evaluated in light of these results. Our method, along with the identified set of contaminated sequences, will be essential for applications that depend on large EST datasets.

Algorithms↗

Serological and genetic characterisation of a unique strain of adenovirus involved in an outbreak of epidemic keratoconjunctivitis.

AIMS: To characterise a novel strain of adenovirus (Ad) type Ad8 (genome type Ad8I) involved in an epidemic keratoconjunctivitis (EKC) outbreak in Hiroshima city using serological testing and sequence analysis of the fibre and hexon gene. METHODS: A neutralisation test (NT) was performed in microtitre plates containing a confluent monolayer of A549 cells using 100 tissue culture infectious doses of virus and type specific antisera. The haemagglutination inhibition test was also carried out in microtitre plates with rat erythrocytes using four haemagglutination units of virus and twofold dilutions of serum. The fibre gene was sequenced by generating overlapping polymerase chain reaction products or by direct sequencing of genomic DNA. Primer selection was based on alignment of the fibre genes of human adenovirus serotypes Ad8, Ad19, Ad37, Ad9, and Ad15 available from Gene Bank. RESULTS: The virus strain was specifically neutralised by anti-Ad8 antibodies, although there was a major crossreaction with anti-Ad9 antibodies. Haemagglutination was equally inhibited by anti-Ad8 and anti-Ad9 antibodies. The predicted amino acid sequences of the hypervariable regions (HVRs) of the Ad8I hexon gene showed higher homology with Ad9 (83.3%) than with Ad8 (62.0%). However, the Ad8I fibre knob was more homologous to Ad8 (94.4%) than to Ad9 (91.6%). CONCLUSIONS: Ad8I is a unique strain of adenovirus because of its lower genomic homology with Ad8, major crossreactivity with Ad9 in NT, and mixed genetic organisation of HVRs of the hexon gene. These factors may have enabled the virus to circumvent acquired immunity, resulting in the outbreak.

Adenovirus Infections, Human↗