Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,297 records · Page 72Linked to original sources

Analysis of low-density lipoprotein receptor gene mutations in a Chinese patient with clinically homozygous familial hypercholesterolemia.

OBJECTIVE: To screen the point mutation of the low-density lipoprotein receptor (LDL-R) gene in Chinese familial hypercholesterolemia (FH) patients, characterize the relationship between the genotype and the phenotype and discuss the molecular pathological mechanism of FH. METHODS: A patient with clinical phenotype of homozygous FH and her parents were investigated for mutations in the promoter and all eighteen exons of the LDL-R gene. Screening was carried out using Touch-down PCR and direct DNA sequencing; multiple alignment analysis by DNASIS 2.5 was used to find base alteration, and the LDL-R gene mutation database was searched to identify the alteration. In addition, the apolipoprotein B gene (apo B) was screened for known mutations (R3500Q) that cause familial defective apo B100 (FDB) by polymerase chain reaction-restriction fragment length polymorphism (PCR-RFLP). RESULTS: Two new heterozygous mutations in exons 4 and 9 of the LDL-R gene were identified in the proband (C122Y and T383I) as well as her parents. Both of the mutations have not been published in the LDL-R gene mutation database. No mutation of apo B100 (R3500Q) was observed. CONCLUSION: Two new mutations (C112Y and T383I) were found in the LDL-R gene, which may result in FH and may be particularly pathogenetic genotypes in Chinese people.

Adult↗

[Cloning and expression of tumor necrosis factor (TNFalpha) cDNA from red seabream pagrus major].

A fragment of TNFalpha cDNA sequence from red seabream was cloned by homology cloning approach with two degenerated primers which were designed based on the conserved regions of other animals' TNF sequences. The sequence was elongated by 3' and 5' RACE to get the full length CDS sequence. This sequence contained 1264 nucleotides that included a 5' UTR of 85 bp, a 3' UTR of 514 bp and an open reading frame (ORF) of 666 bp which could encode 222 amino acids propeptide. In 3' UTR, there were several mRNA instability motifs and three endotoxin-responsive sequences, but the sequence lacked the polyadenylation signal. The deduced peptide had a clear transmembrane domain, a TNFalpha family signature and a TNF2 family profile. The cell attachment sequence and the glycosaminoglycan attachment sites were also found in the sequence. The red seabream TNF sequence shared relatively high similarity with both mammalian TNFalpha and TNFbeta by multiple sequence alignments. Phylogenetic analysis showed that the piscine TNFalpha were located independently in a different branch compared with mammalian TNFalpha and TNFbeta. Based on the primary and secondary structure analysis and gene expression study, we could concluded that the red seabream TNF should be a TNFalpha, not TNFbeta. RT-PCR was used to study TNFalpha transcript expression. 24 h after the red seabream was challenged by Vibrio anguillarum, the RS TNFalpha transcript expression were detected in blood, brain, gill, heart, head kidney, kidney, liver, muscle and spleen. Results showed that TNFalpha mRNA was constitutively expressed in parts of the tissues both in stimulated and unstimulated fish and the expression could be enhanced after the pathogen infection.

Amino Acid Sequence↗

In silico study of breast cancer associated gene 3 using LION Target Engine and other tools.

Sequence analysis of individual targets is an important step in annotation and validation. As a test case, we investigated human breast cancer associated gene 3 (BCA3) with LION Target Engine and with other bioinformatics tools. LION Target Engine confirmed that the BCA3 gene is located on 11p15.4 and that the two most likely splice variants (lacking exon 3 and exons 3 and 5, respectively) exist. Based on our manual curation of sequence data, it is proposed that an additional variant (missing only exon 5) published in a public sequence repository, is a prediction artifact. A significant number of new orthologs were also identified, and these were the basis for a high-quality protein secondary structure prediction. Moreover, our research confirmed several distinct functional domains as described in earlier reports. Sequence conservation from multiple sequence alignments, splice variant identification, secondary structure predictions, and predicted phosphorylation sites suggest that the removal of interaction sites through alternative splicing might play a modulatory role in BCA3. This in silico approach shows the depth and relevance of an analysis that can be accomplished by including a variety of publicly available tools with an integrated and customizable life science informatics platform.

Adaptor Proteins, Signal Transducing↗

Artificial intelligence techniques for bioinformatics.

This review provides an overview of the ways in which techniques from artificial intelligence (AI) can be usefully employed in bioinformatics, both for modelling biological data and for making new discoveries. The paper covers three techniques: symbolic machine learning approaches (nearest neighbour and identification tree techniques), artificial neural networks and genetic algorithms. Each technique is introduced and supported with examples taken from the bioinformatics literature. These examples include folding prediction, viral protease cleavage prediction, classification, multiple sequence alignment and microarray gene expression analysis.

Algorithms↗

[Sequencing of adenovirus type 7 vaccine strain fragment and characterization of the hexon encoding gene].

OBJECTIVE: To complete the full-length sequencing of the human adenovirus type 7 vaccine strain (Ad7v) for novel vector constructing. METHODS: The Ad7v DNA was digested with SalI and the 17.5-68.0 map unit (mu) fragment was cloned and sequenced. The homology of encoding sequence of Ad7v hexon to those of group A,C,D,E,F and other numbers of group B was accomplished with the software CLUSTAL.V. The three-dimensional structure of the Ad7v hexon was predicted with the RasMo12.71. RESULTS: The fragment contains 17,596 bp, part of E2 and late gene L1, L2 and L3 were encoded by this region. Polypeptide encoded by hexon gene lies in L3 region, which is composed of 934 amino acids. Multiple sequence alignment with the other nine known hexon protein sequences suggested that the variable sequences are mainly concentrated on seven regions, namely hypervariable regions (HVRs). The seven HVRs are related to type-specificity and group-specificity. The three-dimensional structure of the Ad7v hexon revealed that the variable regions are located in the I1 and I2 loops of the molecule mostly on the tower of the hexon. CONCLUSION: The full-length genome sequencing of Ad7v was accomplished at last. Since the deduced amino acid sequence of Ad7v hexon was quite different from other adenoviral vectors such as Ad5 and Ad2, this virus can be potentially used for the construction of novel gene delivery vectors to counterpart the immunity to the vectors widely used at present.

Adenovirus E3 Proteins↗

[Cloning and expression of pituitary prolactin gene in Ailuropoda melanoleuca].

The giant panda (Ailuropoda melanoleuca) is an endangered species and indigenous to China. It has been proposed that it has a highly specialized reproductive pattern with low fecundity, but little is known about its basic reproductive biology at molecular level. In this study,the pituitary prolactin (PRL) cDNA of giant panda was amplified by RT-PCR from pituitary total RNA and then cloned, sequenced and submitted to GenBank (GenBank accession No. AY161285). The sequence analysis revealed that the giant panda prolactin cDNA contains a 687-nucleotide open reading frame encoding the prolactin prohormone of 229 amino acid residues. The signal peptide contains 30 amino acid residues and the mature prolactin is composed of 199 amino acid residues. Then the DNA fragment amplified was subcloned into pGEX-4T-1 procaryotic expression plasmid and protein expression was induced by IPTG in Escherichia coil BL21. SDS-PAGE analysis revealed the PRL protein is infusible. The multiple sequence alignments revealed that the homology of giant panda is 95% to cat and pig, 80% - 70% to human, cow and goat, 52% to rat and 45.9% to mouse at the amino acid level. The 64th amino acid of giant panda prolactin is hydrophilic serine instead of hydrophobic proline of cat, goat, and cow or hydrophobic alanine of human.

Amino Acid Sequence↗

A search tool for identification and analysis of conserved sequence patterns in Saccharomyces spp. orthologous promoter.

We describe a web-based resource to identify, search and analyze sequence patterns conserved in the multiple sequence alignments of orthologous promoters from closely related / distant Saccharomyces spp. The webtool interfaces with a database where conserved sequence patterns (greater than 4 bp) have been previously extracted from genome-wide promoter alignments, allowing one to carry out user-defined genome-wide searches for conserved sequences to assist in the discovery of novel promoter elements based on comparative genomics. The web-based server can be accessed at http://www2.imtech.res.in/ anand/sacch_prom_pat.html.

Base Sequence↗

Molecular cloning and characterization of a glutathione S-transferase encoding gene from Opisthorchis viverrini.

An adult stage Opisthorchis viverrini cDNA library was constructed and screened for abundant transcripts. One of the isolated cDNAs was found by sequence comparison to encode a glutathione S-transferase (GST) and was further analyzed for RNA expression, encoded protein function, tissue distribution and cross-reactivity of the encoded protein with other trematode protein counterparts. The cDNA has a size of 893 bp and encodes a GST of 213 amino acids length (OV28GST). The most closely-related GST of OV28GST among those published for trematodes is a 28 kDa GST of Clonorchis sinensis as shown by multiple sequence alignment and phylogenetic analysis. Northern analysis of total RNA with a gene-specific probe revealed a 900 nucleotide OV28GST transcriptional product in the adult parasite. Through RNA in situ hybridization OV28GST RNA was detected in the parenchymal cells of adult parasites. This result was confirmed by immunolocalization of OV28GST with an antiserum generated in a mouse against bacterially-produced recombinant OV28GST. Both, purified recombinant and purified native OV28GST were resolved as 28 kDa proteins by SDS-PAGE. Using the anti-recOV28GST antiserum, no or only weak cross-reactivity was observed in an immunoblot of crude worm extracts against the GSTs of Schistosoma mansoni, S. japonicum, S. mekongi, Eurytrema spp. and Fasciola gigantica. The enzyme activity of the purified recombinant OV28GST was verified by a standard 1-chloro-2, 4-dinitrobenzene (CDNB) based activity assay. The present results of our molecular analysis of OV28GST should be helpful in the ongoing development of diagnostic applications for opisthorchiasis viverrini.

Amino Acid Sequence↗

[Isolation and expression profiling of the Pto-like gene SsPto from Solanum surattense].

A novel Pto-like gene (designated as SsPto) is cloned from yellow-fruit nightshade (Solanum surattense). The full-length cDNA of SsPto is 1331 bp long with an open reading frame of 960 bp encoding a polypeptide of 320 amino acid residues. The deduced SsPto protein has a calculated molecular weight of 36.21 kDa with an isoelectric point of 6.18. Multiple sequence alignment shows that SsPto protein shares 71.4% and 71.6% identities to Pto proteins from Lycopersicon pimpinellifolium and L. hirsutum respectively. Genomic Southern blot analysis indicates the presence of a small family of SsPto in the S. surattense genome. SsPto is found to be constitutively expressed in the S. surattense plant with the highest expression in stems. However, under induction by TMV for 6 days, SsPto expresses the highest in roots. Further expression analysis reveals that the signaling components of defense/stress pathways, such as methyl jasmonate (MeJA), salicylic acid (SA), gibberellic acid (GA3) and hydrogen peroxide (H2O2), up-regulate the SsPto transcript levels over the control. Cold treatment, nevertheless, has no significant effect on SsPto expression whereas SsPto expression is down-regulated by dark treatment. Our findings suggest that this novel stress- and pathogen-inducible SsPto from S. surattense may participate not only in the defense/stress responsive pathways, but also in diverse processes of plant's growth and development.

Amino Acid Sequence↗

New strategy to detect single nucleotide polymorphisms.

A great effort has been made to identify and map a large set of single nucleotide polymorphisms. The goal is to determine human DNA variants that contribute most significantly to population variation in each trait. Different algorithms and software packages, such as PolyBayes and PolyPhred, have been developed to address this problem. We present strategies to detect single nucleotide polymorphisms, using chromatogram analysis and consensi of multiple aligned sequences. The algorithms were tested using HIV datasets, and the results were compared with those produced by PolyBayes and PolyPhred using the same dataset. Our algorithms produced significantly better results than these two software packages.

Algorithms↗

SSToSS--sequence-structural templates of single-member superfamilies.

The presence of sequence homologues and the availability of structural information of proteins enable better understanding of the biological function of a protein family. A majority of entries in protein structural databank are single member superfamilies for which it is hard to derive motifs due to the paucity of structural homologues. Important conserved segments for these superfamilies have been identified and compiled into a database, SSToSS (Sequence Structural Templates of Single member Superfamily). Conserved regions, recognized by permitted amino acid exchanges, are mapped on the structure and various structural features (solvent accessibility, secondary structure content, hydrogen bonding and residue packing) are examined. These conserved segments with high structural feature content are projected as sequence-structural templates for the particular superfamily member. Interactive three-dimensional displays of the templates in three-dimensional structure (in Chime and RASMOL) are provided for better understanding and visualization. In SSToSS database, we also provide the application of sequence-structural templates in three different areas: multiple-motif based sequence search, multiple sequence alignment and homology modeling. In each case, the inclusion of the sequence-structural templates can give rise to sensitive and accurate results. This enables the inclusion of singletons to provide added value to the recognition of additional members, comparative modeling and in designing experiments.

Amino Acid Motifs↗

[Rapid detection of Pseudomonas aeruginosa by the fluorescence quantitative PCR assay targeting 16S rDNA].

The 16S rDNA specific primers were designed for rapid detection of Pseudomonas aeruginosa (PA) by the fluorescence quantitative PCR (FQ-PCR) assay, based upon multiple sequence alignment and phylogenetic tree analysis of the 16S rDNAs of over 20 bacteria. After extraction of PA genomic DNA, the target 16S rDNA fragment was amplified by PCR with specific primers, and used to construct recombinant pMDT-Pfr plasmid, the dilution gradients of which were subjected to the standard quantitation curve in FQ-PCR assay. Different concentrations of PA genomic DNA were detected by FQ-PCR in a 20microL of reaction system with SYBR Green I. At the same time, various genomic DNAs of Staphylococcus aureus, Salmonella typhi, Shigella flexneri, Proteus vulgaris, Staphylococcus epidermidis, Escherichia coli, and Mycobacterium tuberculosis were used as negative controls to confirm specificity of the FQ-PCR detection assay. Results demonstrated that the predicted amplified product of designed primers was of high homology only with PA 16S rDNA, and that sensitivity of the FQ-PCR assay was of 3.6pg/microL of bacterial DNA or (2.1 x 10(3) +/- 3.1 x 10(2)) copies/microL of 16S rDNA, accompanied with high specificity, and that the whole detection process including DNA extraction could be completed in about two hours. In contrast to traditional culture method, the FQ-PCR assay targeting 16S rDNA gene can be used to detect PA rapidly, which exhibits perfect application prospect in future.

Base Sequence↗

The megaprior heuristic for discovering protein sequence patterns.

Several computer algorithms for discovering patterns in groups of protein sequences are in use that are based on fitting the parameters of a statistical model to a group of related sequences. These include hidden Markov model (HMM) algorithms for multiple sequence alignment, and the MEME and Gibbs sampler algorithms for discovering motifs. These algorithms are sometimes prone to producing models that are incorrect because two or more patients have been combined. The statistical model produced in this situation is a convex combination (weighted average) of two or more different models. This paper presents a solution to the problem of convex combinations in the form of a heuristic based on using extremely low variance Dirichlet mixture priors as part of the statistical model. This heuristic, which we call the megaprior heuristic, increase the strength (i.e., decreases the variance) of the prior in proportion to the size of the sequence dataset. This causes each column in the final model to strongly resemble the mean of a single component of the prior, regardless of the size of the dataset. We describe the cause of the convex combination problem, analyze it mathematically, motivate and describe the implementation of the megaprior heuristic, and show how it can effectively eliminate the problem of convex combinations in protein sequence pattern discovery.

Algorithms↗

Prediction of the structure of the replication initiator protein DnaA.

The secondary structure of DnaA protein and its interaction with DNA and ribonucleotides has been predicted using biochemical, biophysical techniques, and prediction methods based on multiple-sequence alignment and neural networks. The core of all proteins from the DnaA family consists of an "open twisted alpha/beta structure," containing five alpha-helices alternating with five beta-strands. In our proposed structural model the interior of the core is formed by a parallel beta-sheet, whereas the alpha-helices are arranged on the surface of the core. The ATP-binding motif is located within the core, in a loop region following the first beta-strand. The N-terminal domain (80 aa) is composed of two alpha-helices, the first of which contains a potential leucine zipper motif for mediating protein-protein interaction, followed by a beta-strand and an additional alpha-helix. The N-terminal domain and the alpha/beta core region of DnaA are connected by a variable loop (45-70 aa); major parts of the loop region can be deleted without loss of protein activity. The C-terminal DNA-binding domain (94 aa) is mostly alpha-helical and contains a potential helix-loop-helix motif. DnaA protein does not dimerize in solution; instead, the two longest C-terminal alpha-helices could interact with each other, forming an internal "coiled coil" and exposing highly basic residues of a small loop region on the surface, probably responsible for DNA backbone contacts.

Amino Acid Sequence↗

Structure and genomic organization of a second class of immunoglobulin light chain genes in the channel catfish.

Earlier studies distinguished two classes of catfish light (L) chain (designated F and G). The cDNA structure and genomic organization of G L chain gene clusters has also been characterized previously. In this study, full length cDNA encoding F L chain was derived using PCR strategies based on the determined amino-terminal protein sequence. The encoded V region is readily delineated into framework regions (FR) and complementarity-determining regions (CDR). Multiple sequence alignments indicate that the F V(L) is closely related to kappa gene families. The F C(L) cannot be generally classified but it is structurally distinct from the C(L) regions of G: the amino acid sequence similarity is <35%. cDNA sequences representing processed sterile F transcripts of different loci were identified. Each sequence begins within the J(L) recombination signal sequence and extends downstream through the I(L)-C(L) segments. Genomic blots hybridized with C(L) probes indicate that there are at least 50 different C(L) segments. Based upon V(L) hybridization studies, different families of V(L) segments appear to be associated with closely related F C(L) segments. In characterized genomic clones, F gene segments are arranged in closely linked clusters with single copies of V(L), J(L), and C(L) segments within each cluster. The V(L) segments are located in opposite transcriptional polarity relative to the J(L) and C(L) segments, which indicates that V(L) segments rearrange by inversion. These combined studies establish that two structurally distinct classes of L chains are present in teleost fish and that both of the L chain classes evolved within a common organizational pattern of clustered segmental genes.

Amino Acid Sequence↗

A model for the nucleotide-binding domains of ABC transporters based on the large domain of aspartate aminotransferase.

ABC transporters are a large superfamily of integral membrane proteins involved inATP-dependent transport across biological membranes. Members of this superfamily play roles in a number of phenomena of biomedical interest, including cystic fibrosis (CFTR) and multidrug resistance (P-glycoprotein, MRP). Most ABC transporters are predicted to consist of four domains, two membrane-spanning domains and two cytoplasmic domains. The latter contain conserved nucleotide-binding motifs. Attempts to determine the structure of ABC transporters and of their separate domains are in progress but have not yet been successful. To aid structure determination and possibly learn more about the domain boundaries, we set out to model nucleotide-binding domains (NBDs) of ABC transporters based on a known structure. Previous attempts to predict the 3D structure of NBDs were based solely on sequence similarity with known nucleotide-binding folds. We have analyzed the sequences of a number of nucleotide-binding domains with the algorithm THREADER, developed by D.T. Jones, and a possible fold was found in the structure of aspartate aminotransferase. We present a model for the N-terminal NBD of CFTR, based on the large domain of the A chain of aspartate aminotransferase. The model is refined using multiple sequence alignment, secondary structure prediction, and 3D-1D profiles. Our model seems to be in good agreement with known properties of nucleotide-binding domains and has some appealing characteristics compared with the previous models.

ATP-Binding Cassette Transporters↗

Isolation and characterization of a skate retinal GABA transporter cDNA.

PURPOSE: The inhibitory neurotransmitter gamma-aminobutyric acid (GABA) is believed to play a crucial role in the processing of information within the vertebrate retina. Extracellular concentrations of GABA are thought to be tightly regulated by carrier-mediated transport proteins in neurons and glial cells. The purpose of this work was to isolate the gene that encodes one of these transport proteins in the skate retina. METHODS: cDNA clones were isolated from a skate retinal cDNA library using a mouse retinal GABA transporter (GAT1) cDNA as a probe. The PCR technique was used to fill sequence gaps, and 5' and 3' RACE were employed to amplify the 5' and 3' untranslated regions. The amplified fragments were subcloned into a T-vector. Blots containing RNA from 10 different tissues were probed to determine the size of the transcript and the tissue distribution. RESULTS: Sequence analysis revealed that the skate retinal GABA transporter cDNA shared 72% identity with the mouse GABA transporter-1 at the DNA level and 80% identity at the amino acid level. Multiple sequence alignments showed that our sequence is closest to the Torpedo GABA transporter-1. Two transcripts, 4.5 and 7 kb, were detected in retina and possibly brain by RNA blot analysis. Fourteen introns were detected in the skate GABA transporter gene. CONCLUSIONS: We successfully isolated a full length GABA transporter cDNA from the retina of the skate. The size of the full length sequence of the skate retinal GABA transporter is in agreement with the size of the smaller transcript detected on RNA blots. The larger transcript observed on the RNA blot may be the result of either alternative splicing or utilization of a downstream poly A signal.

Animals↗

AAA+: A class of chaperone-like ATPases associated with the assembly, operation, and disassembly of protein complexes.

Using a combination of computer methods for iterative database searches and multiple sequence alignment, we show that protein sequences related to the AAA family of ATPases are far more prevalent than reported previously. Among these are regulatory components of Lon and Clp proteases, proteins involved in DNA replication, recombination, and restriction (including subunits of the origin recognition complex, replication factor C proteins, MCM DNA-licensing factors and the bacterial DnaA, RuvB, and McrB proteins), prokaryotic NtrC-related transcription regulators, the Bacillus sporulation protein SpoVJ, Mg2+, and Co2+ chelatases, the Halobacterium GvpN gas vesicle synthesis protein, dynein motor proteins, TorsinA, and Rubisco activase. Alignment of these sequences, in light of the structures of the clamp loader delta' subunit of Escherichia coli DNA polymerase III and the hexamerization component of N-ethylmaleimide-sensitive fusion protein, provides structural and mechanistic insights into these proteins, collectively designated the AAA+ class. Whole-genome analysis indicates that this class is ancient and has undergone considerable functional divergence prior to the emergence of the major divisions of life. These proteins often perform chaperone-like functions that assist in the assembly, operation, or disassembly of protein complexes. The hexameric architecture often associated with this class can provide a hole through which DNA or RNA can be thread; this may be important for assembly or remodeling of DNA-protein complexes.

Adenosine Triphosphatases↗