Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 865 records · Page 48Linked to original sources

Comparison of two b1 alleles from within the A mating-type of the basidiomycete Coprinus cinereus.

We cloned and sequenced the b1 specificity gene, b1-2, from the A43 mating-type locus of the basidiomycete Coprinus cinereus, to compare its molecular structure to a previously published allele. The b1-2 gene was identified and isolated using a transformation assay for A-activity. The nucleotide (nt) sequence was determined and compared to the published sequence for the b1 specificity gene of the A42 mating-type locus. Both genes map to the same physical location within the A mating-type locus and conserved structural organization is observed at both the genomic and the protein level. Sequence alignments show that the two alleles share 73% overall nt sequence identity for the open reading frames (ORFs) and 68% overall amino acid (aa) sequence identity for the deduced polypeptides. Allowing for conservative substitutions, the overall aa sequence similarity is 79%. Comparison of the deduced aa sequences reveals several conserved structural motifs, including a DNA-binding homeodomain, putative bipartite nuclear localization signal sequences, and four predicted dimerization motifs. Regions rich in Pro and hydroxylated aa (Ser and Thr) are also common to both alleles. Sequence similarity varies greatly along the length of the gene at both the nt and aa levels. In general, similarity increases progressively from the N- to the C-terminal end with variable patterns of similarity observed for individual exons and the predicted motifs they encode. The low sequence similarity observed in the N terminus suggests that variability in this region may be involved in non-self recognition. Differing levels of positional and compositional constraint are apparent for different functional domains.

Alleles↗

Transcriptomic changes in the gut mucosa of fasting northern elephant seal pups reveal immune modulation during early microbiome establishment.

Fasting is an integral component of the life-history of many species. Following abrupt weaning, northern elephant seal pups (Mirounga angustirostris) undergo an extended post-weaning fast of approximately 60 days. During this period, enteric bacterial diversity increases, suggesting that host immune regulation may facilitate the establishment of microbial communities. However, the molecular processes occurring within the intestinal mucosa during this transition remain poorly understood. To investigate these mechanisms, we characterized transcriptional changes in the enteric mucosa of male and female northern elephant seal pups sampled at weaning and after one month of fasting. Total RNA isolated from rectal swabs was sequenced and aligned to the Mirounga angustirostris reference genome. Differential gene expression and gene set enrichment analyses were used to identify genes and pathways associated with fasting and sex-specific responses. Fasting was accompanied primarily by transcriptional downregulation, including genes involved in antimicrobial defense, inflammation, protein turnover, and epithelial remodeling. In contrast, several genes associated with B-cell activity and immune recognition were upregulated. Gene Set Enrichment Analysis revealed coordinated activation of immune-regulatory pathways indicating dynamic modulation of intestinal immunity rather than generalized immune suppression. Pronounced sex-specific differences were also observed. Male pups exhibited transcriptional patterns consistent with enhanced immune tolerance, whereas females showed broader immune-pathway activation, including enrichment of pro-inflammatory and stress-response pathways. Several non-coding RNAs also displayed sex-specific changes in expression. Together, these findings suggest that fasting induces transcriptional remodeling of the gut and may contribute to immune regulation during a critical period of microbiome establishment in northern elephant seal pups.

Animals↗

Cloning, gene organization and identification of an alternative splicing process in lecithin:retinol acyltransferase cDNA from human liver.

Lecithin:retinol acyltransferase (LRAT) catalyzes the synthesis of retinyl esters in many tissues and is crucial for the transport and intracellular storage of vitamin A. LRAT expression is highly regulated in the liver. In this study, we have cloned and sequenced the full-length LRAT mRNA from human liver and identified its 5'- and 3'-ends. Full-length LRAT mRNA comprises 5023 nt with a predicted ORF of 230 amino acids, a short 5'UTR, and a relatively long 3'UTR of 4 kb containing several polyadenylation signals and AU-rich regions. Based on alignment of this mRNA with human genomic DNA in the GenBank database, the human LRAT gene spans about 9.1 kbp and consists of two exons and a relatively long 4-kbp intron. Further analysis of normal liver revealed a minor alternative splicing variant which lacks a 103 nt polynucleotide contained in the 5'UTR of the full-length LRAT transcript. This variant predicts that the LRAT gene is organized into three exons and two introns, as reported for LRAT cloned from retinal pigment epithelium (RPE) cells. These two LRAT mRNA variants are also present in testis, which is known to express LRAT and contain retinyl esters. Major and minor transcription start sites for human liver LRAT mRNA were identified and the sequence of the upstream proximal promoter region was retrieved from the GenBank database and physically analyzed for the presence of putative cis-acting elements essential for basal transcription. This region contains a TATA box, CCAAT box and Sp1 site, which are apparently conserved in mouse and rat LRAT genes. Our results provide evidence that multiple LRAT mRNA transcripts, which are expressed in a tissue-specific manner, may result from several mechanisms including differential splicing of the 5'UTR region and the use of multiple polyadenylation signals in the 3'UTR.

3' Untranslated Regions↗

SayaMatcher: genome scale organization and systematic analysis of nuclear receptor response elements.

The availability of genome sequences enables us to make experiments in silico to find biological features encoded in the genome. In order to study gene expression regulation network by nuclear receptors that function as ligand-activated transcription, nuclear receptor response elements (NREs) in genomes were computationally explored by integration of computational prediction and experimental data. Dealing with expansion of available genome sequences and continuous update of these genomic sequences, the computation needed for these searches is organized to form a 'pipeline', called SayaMatcher. The information of genomic position for those NREs is served using the Distributed Annotation System (DAS) to be shown in Ensembl genome browser. Utilizing the SayaMatcher system, binding activity in vivo was studied for androgen response elements in human.

Animals↗

Comparative proteomics of the Mycobacterium leprae binding protein myelin P0: its implication in leprosy and other neurodegenerative diseases.

Mycobacterium leprae, the causative agent of leprosy invades Schwann cells of the peripheral nerves leading to nerve damage and disfigurement, which is the hallmark of the disease. Wet experiments have shown that M. leprae binds to a major peripheral nerve protein, the myelin P zero (P0). This protein is specific to peripheral nerve and may be important in the initial step of M. leprae binding and invasion of Schwann cells which is the feature of leprosy. Though the receptors on Schawann cells, cytokines, chemokines and antibodies to M. leprae have been identified the molecular mechanism of nerve damage and neurodegeneration is not clearly defined. Recently pathogen and host protein/nucleotide sequence similarities (molecular mimicry) have been implicated in neurodegenerative diseases. The approach of the present study is to utilise bioinformatic tools to understand leprosy nerve damage by carrying out sequence and structural similarity searches of myelin P0 with leproma and other genomic database. Since myelin P0 is unique to peripheral nerve, its sequence and structural similarities in other neuropathogens have also been noted. Comparison of myelin P0 with the M. leprae proteins revealed two characterised proteins, Ferrodoxin NADP reductase and a conserved membrane protein, which showed similarity to the query sequence. Comparison with the entire genomic database (www.ncbi.nlm.nih.gov) by basic local alignment search tool for proteins (BLASTP) and fold classification of structure-structure alignment of proteins (FSSP) searches revealed that myelin P0 had sequence/structural similarities to the poliovirus receptor, coxsackie-adenovirus receptor, anthrax protective antigen, diphtheria toxin, herpes simplex virus, HIV gag-1 peptide, and gp120 among others. These proteins are known to be associated directly or indirectly with neruodegeneration. Sequence and structural similarities to the immunoglobin regions of myelin P0 could have implications in host-pathogen interactions, as it has homophilic adhesive properties. Although these observed similarities are not highly significant in their percentage identity, they could be functionally important in molecular mimicry, receptor binding and cell signaling events involved in neurodegeneration.

Amino Acid Sequence↗

Protein domain analysis in the era of complete genomes.

Domains present one of the most useful levels at which to understand protein function, and domain family-based analysis has had a profound impact on the study of individual proteins. Protein domain discovery has been progressing steadily over the past 30 years. What are the realistically achievable goals of sequence-based domain analysis, and how far off are they for the sequences encoded in eukaryotic genomes? Here we address some of the issues involved in better coverage of sequence-based domain annotation, and the integration of these results within the wider context of genomes, structures and function.

Amino Acid Sequence↗

Characterization of two putative histone deacetylase genes from Aspergillus nidulans.

In eukaryotic organisms, acetylation of core histones plays a key role in the regulation of transcription. Multiple histone acetyltransferases (HATs) and histone deacetylases (HDACs) maintain a dynamic equilibrium of histone acetylation. The latter form a highly conserved protein family in many eukaryotic species. In this paper, we report the cloning and sequencing of two putative histone deacetylase genes (rpdA, hosA) of Aspergillus nidulans, which are the first to be analyzed from filamentous fungi. Hybridization with a chromosome-specific cosmid library of A. nidulans allowed the localization of rpdA to chromosome III and hosA to chromosome II, respectively. PCR analyses and Southern hybridization experiments revealed that no further members of the RPD3 family are present in the genome of the fungus. Although sequence alignment displays significant amino acid similarity to other eukaryotic RPD3-type deacetylases, the deduced RPDA sequence reveals an unusual 200-amino acid extension at the C-terminus. Expression of both genes was determined by RNA blot analysis. Treatment of the cells with trichostatin A (TSA), a potent inhibitor of HDACs, was found to stimulate expression of rpdA of A. nidulans.

Amino Acid Sequence↗

Identification and localization of two mouse phosphomannomutase genes, Pmm1 and Pmm2.

Phosphomannomutases catalyze the reversible conversion of mannose 6-phosphate to mannose 1-phosphate. In humans, two different isozymes have recently been identified, PMM1 and PMM2. We have previously shown that mutations in the PMM2 gene cause the most frequent type of the congenital disorders of glycosylation, CDG-Ia. Here, we present data on the two mouse orthologous genes, Pmm1 and Pmm2. The chromosomal localization of the two mouse genes has been determined. We also present the gene structure and the exon-intron organization of Pmm1 and Pmm2. Pmm1 maps to mouse chromosome 15, Pmm2 to chromosome 16. These chromosomal regions are syntenic with regions on human chromosomes 22 and 16, respectively. The Pmm1 gene is composed of eight exons and spans approximately 9.5 kb. The genomic structure is extremely well conserved between the human and mouse gene. The Pmm2 gene consists of eight exons and spans a larger genomic region ( approximately 20 kb). An alignment of the human and mouse protein sequences confirms the conservation among this family of phosphomannomutases. The two mouse genes are expressed in many tissues, but the expression pattern is slightly different between Pmm1 and Pmm2. The most striking difference is the high expression of Pmm1 in brain tissue, whereas Pmm2 is only weakly expressed in this tissue.

Amino Acid Sequence↗

Protein tyrosine phosphatases: counting the trees in the forest.

The recent identification of many different protein tyrosine phosphatases (PTPs) has led to the recognition that these enzymes match protein tyrosine kinases (PTKs) in importance for intracellular signalling. The total number of PTPs encoded by the mammalian genome has been estimated at between 500 and approx. 2000. These estimates are imprecise due to the large number of sequence database entries that represent different splice forms, or duplicates of the same PTP sequence. A careful analysis of these entries, grouped by identical catalytic domain shows that no more than 48 full-length PTP sequences are currently known, and that their total number in the human genome may not exceed 100. An alignment of all catalytic domains also suggests that during evolution intragenic catalytic domain duplication, as seen in most membrane-bound PTPs, preceded gene duplication.

Amino Acid Sequence↗

Overview of structural genomics: from structure to function.

The unprecedented increase in the number of new protein sequences arising from genomics and proteomics highlights directly the need for methods to rapidly and reliably determine the molecular and cellular functions of these proteins. One such approach, structural genomics, aims to delineate the total repertoire of protein folds, thereby providing three-dimensional portraits for all proteins in a living organism and to infer molecular functions of the proteins. The goal of obtaining protein structures on a genomic scale has motivated the development of high-throughput technologies for macromolecular structure determination, which have begun to produce structures at a greater rate than previously possible. These new structures have revealed many unexpected functional and evolution relationships that were hidden at the sequence level.

Amino Acid Sequence↗

Arenavirus phylogeny: a new insight.

Arenaviridae is a worldwide distributed family, of enveloped, single stranded, RNA viruses. The arenaviruses were divided in two major groups (Old World and New World), based on serological properties and genetic data, as well as the geographic distribution. In this study the phylogenetic relationship among the members of the Arenaviridae was examined, using the reported genomic sequences. The comparison of the aligned nucleotide sequences of the S RNA and the predicted amino acid sequences of the GPC and N proteins, together with the phylogenetic analysis, strongly suggest a possible kinship of Pichindé and Oliveros viruses, with the Old World arenavirus group. This analysis points at the evolutive relationships between the arenaviruses of the Americas and can be used to evaluate the different hypotheses about their origin.

Arenavirus↗

Alternative splicing in the brain of mice and rats generates transferrin transcripts lacking, as in humans, the signal peptide sequence.

Transferrin (Tf), the iron-transport protein of vertebrate serum, is mainly synthesized in hepatocytes but is also found in other cell-types including oligodendrocytes. Our laboratory has characterized in a human oligodendrial cell line the presence of a new Tf transcript containing an alternative exon 1b replacing the classical exon 1 and conducting to the elimination of the signal peptide sequence. In this manuscript, we show by RT-PCR and 5'-RACE experiments that alternative transcripts also exist in mouse and rat and are found in brain mRNA preparations. Mouse alternative first exon is homologous to human exon 1b while rat Tf gene was found to use a new first exon named 1c. In all species, the alternative transcript does not contain the signal peptide sequence and possibly encode for a Tf protein devoid of signal peptide showing that this phenomenon is not restricted to human gene. We also present genomic sequence data from the previously unknown 5' genomic rat region, which allowed the alignment of the alternative exons 1 in the three species.

Alternative Splicing↗

Characterization of the human properdin gene.

A cosmid clone containing the complete coding sequence of the human properdin gene has been characterized. The gene is located at one end of the approximately 40 kb cosmid insert and approximately 8.2 kb of the sequence data have been obtained from this region. Two discrepancies with the published cDNA sequence [Nolan, Schwaeble, Kaluz, Dierich & Reid (1991) Eur. J. Immunol. 21, 771-776] have been resolved. Properdin has previously been described as a modular protein, with the majority of its sequence composed of six tandem repeats of a sequence motif of approximately 60 amino acids which is related to the type-I repeat sequence (TSR), initially described in thrombospondin [Lawler & Hynes (1986) J. Cell Biol. 103, 1635-1648; Goundis & Reid (1988), Nature (London) 335, 82-85]. Analysis of the genomic sequence data indicates that the human properdin gene is organized into ten exons which span approximately 6 kb of the genome. TSRs 2-5 are coded for by discrete, symmetrical exons (phase 1-1), which supports the hypothesis that modular proteins evolved by a process involving exon shuffling. TSR1 is also coded for by a discrete exon, but the boundaries are asymmetrical (phase 2-1). The sequence coding for the sixth TSR is split across the final two exons of the gene with the first 38 amino acids of the repeat coded for by an asymmetric exon (phase 1-2). This split at the genomic level has been shown, by alignment analysis, to be reflected at the protein level with the division of repeat 6 into TSR-like and TSR-unlike sequences.

Amino Acid Sequence↗

Keratin K6irs is specific to the inner root sheath of hair follicles in mice and humans.

BACKGROUND: Keratins are a multigene family of intermediate filament proteins that are differentially expressed in specific epithelial tissues. To date, no type II keratins specific for the inner root sheath of the human hair follicle have been identified. OBJECTIVES: To characterize a novel type II keratin in mice and humans. METHODS: Gene sequences were aligned and compared by BLAST analysis. Genomic DNA and mRNA sequences were amplified by polymerase chain reaction (PCR) and confirmed by direct sequencing. Gene expression was analysed by reverse transcription (RT)-PCR in mouse and human tissues. A rabbit polyclonal antiserum was raised against a C-terminal peptide derived from the mouse K6irs protein. Protein expression in murine tissues was examined by immunoblotting and immunofluorescence. RESULTS: Analysis of human expressed sequence tag (EST) data generated by the Human Genome Project revealed a fragment of a novel cytokeratin mRNA with characteristic amino acid substitutions in the 2B domain. No further human ESTs were found in the database; however, the complete human gene was identified in the draft genome sequence and several mouse ESTs were identified, allowing assembly of the murine mRNA. Both species' mRNA sequences and the human gene were confirmed experimentally by PCR and direct sequencing. The human gene spans more than 16 kb of genomic DNA and is located in the type II keratin cluster on chromosome 12q. A comprehensive immunohistochemical survey of expression in the adult mouse by immunofluorescence revealed that this novel keratin is expressed only in the inner root sheath of the hair follicle. Immunoblotting of murine epidermal keratin extracts revealed that this protein is specific to the anagen phase of the hair cycle, as one would expect of an inner root sheath marker. In humans, expression of this keratin was confirmed by RT-PCR using mRNA derived from plucked anagen hairs and epidermal biopsy material. By this means, strong expression was detected in human hair follicles from scalp and eyebrow. Expression was also readily detected in human palmoplantar epidermis; however, no expression was detected in face skin despite the presence of fine hairs histologically. CONCLUSIONS: This new keratin, designated K6irs, is a valuable histological marker for the inner root sheath of hair follicles in mice and humans. In addition, this keratin represents a new candidate gene for inherited structural hair defects such as loose anagen syndrome.

Amino Acid Sequence↗

Genomic sequence, structural organization, molecular evolution, and aberrant rearrangement of promyelocytic leukemia zinc finger gene.

The promyelocytic leukemia zinc finger gene (PLZF) is involved in chromosomal translocation t(11;17) associated with acute promyelocytic leukemia. In this work, a 201-kilobase genomic DNA region containing the entire PLZF gene was sequenced. Repeated elements account for 19.83%, and no obvious coding information other than PLZF is present over this region. PLZF contains six exons and five introns, and the exon organization corresponds well with protein domains. There are at least four alternative splicings (AS-I, -II, -III, and -IV) within exon 1. AS-I could be detected in most tissues tested whereas AS-II, -III, and -IV were present in the stomach, testis, and heart, respectively. Although splicing donor and acceptor signals at exon-intron boundaries for AS-I and exons 1-6 were classical (gt-ag), AS-II, -III, and -IV had atypical splicing sites. These alternative splicings, nevertheless, maintained the ORF and may encode isoforms with absence of important functional domains. In mRNA species without AS-I, there is a relatively long 5' UTR of 6.0 kilobases. A TATA box and several transcription factor binding sites were found in the putative promoter region upstream of the transcription start site. PLZF is a well conserved gene from Caenorhabditis elegans to human. PLZF paralogous sequences are found in human genome. The presence of two MLL/PLZF-like alignments on human chromosome 11q23 and 19 suggests a syntenic replication during evolution. The chromosomal breakpoints and joining sites in the index acute promyelocytic leukemia case with t(11;17) also were characterized, which suggests the involvement of DNA damage-repair mechanism.

Alternative Splicing↗

Critical aspartic acid residues in pseudouridine synthases.

The pseudouridine synthases catalyze the isomerization of uridine to pseudouridine at particular positions in certain RNA molecules. Genomic data base searches and sequence alignments using the first four identified pseudouridine synthases led Koonin (Koonin, E. V. (1996) Nucleic Acids Res. 24, 2411-2415) and, independently, Santi and co-workers (Gustafsson, C., Reid, R., Greene, P. J., and Santi, D. V. (1996) Nucleic Acids Res. 24, 3756-3762) to group this class of enzyme into four families, which display no statistically significant global sequence similarity to each other. Upon further scrutiny (Huang, H. L., Pookanjanatavip, M., Gu, X. G., and Santi, D. V. (1998) Biochemistry 37, 344-351), the Santi group discovered that a single aspartic acid residue is the only amino acid present in all of the aligned sequences; they then demonstrated that this aspartic acid residue is catalytically essential in one pseudouridine synthase. To test the functional significance of the sequence alignments in light of the global dissimilarity between the pseudouridine synthase families, we changed the aspartic acid residue in representatives of two additional families to both alanine and cysteine: the mutant enzymes are catalytically inactive but retain the ability to bind tRNA substrate. We have also verified that the mutant enzymes do not release uracil from the substrate at a rate significant relative to turnover by the wild-type pseudouridine synthases. Our results clearly show that the aligned aspartic acid residue is critical for the catalytic activity of pseudouridine synthases from two additional families of these enzymes, supporting the predictive power of the sequence alignments and suggesting that the sequence motif containing the aligned aspartic acid residue might be a prerequisite for pseudouridine synthase function.

Amino Acid Sequence↗

The fission yeast protein Ker1p is an ortholog of RNA polymerase I subunit A14 in Saccharomyces cerevisiae and is required for stable association of Rrn3p and RPA21 in RNA polymerase I.

A heterodimer formed by the A14 and A43 subunits of RNA polymerase (pol) I in Saccharomyces cerevisiae is proposed to correspond to the Rpb4/Rpb7 and C17/C25 heterodimers in pol II and pol III, respectively, and to play a role(s) in the recruitment of pol I to the promoter. However, the question of whether the A14/A43 heterodimer is conserved in eukaryotes other than S. cerevisiae remains unanswered, although both Rpb4/Rpb7 and C17/C25 are conserved from yeast to human. To address this question, we have isolated a Schizosaccharomyces pombe gene named ker1+ using a yeast two-hybrid system, including rpa21+, which encodes an ortholog of A43, as bait. Although no homolog of A14 has previously been found in the S. pombe genome, functional characterization of Ker1p and alignment of Ker1p and A14 showed that Ker1p is an ortholog of A14. Disruption of ker1+ resulted in temperature-sensitive growth, and the temperature-sensitive deficit of ker1delta was suppressed by overexpression of either rpa21+ or rrn3+, which encodes the rDNA transcription factor Rrn3p, suggesting that Ker1p is involved in stabilizing the association of RPA21 and Rrn3p in pol I. We also found that Ker1p dissociated from pol I in post-log-phase cells, suggesting that Ker1p is involved in growth-dependent regulation of rDNA transcription.

Amino Acid Sequence↗