Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Molecular Sequence Annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Functional analysis of a novel nonsense PPP1R12A variant in a Chinese family with infantile epilepsy.

BACKGROUND: Defects in PPP1R12A can lead to genitourinary and/or brain malformation syndrome (GUBS). GUBS is primarily characterized by neurological or genitourinary system abnormalities, but a few reported cases are associated with neonatal seizures. Here, we report a case of a female newborn with neonatal seizures caused by a novel variant in PPP1R12A, aiming to enhance the clinical and variant data of genetic factors related to epilepsy in early life. METHODS: Whole-exome and Sanger sequencing were used for familial variant assessment, and bioinformatics was employed to annotate the variant. A structural model of the mutant protein was simulated using molecular dynamics (MD), and the free binding energy between PPP1R12A and PPP1CB was analyzed. A mutant plasmid was constructed, and mutant protein expression was analyzed using western blotting (WB), and the interaction between the mutant and PPP1CB proteins using co-immunoprecipitation (Co-IP) experiments. RESULTS: The patient experienced tonic-clonic seizures on the second day after birth. Genetic testing revealed a heterozygous variant in PPP1R12A, NM_002480.3:c.2533 C > T (p.Arg845Ter). Both parents had the wild-type gene. MD suggested that loss of the C-terminal structure in the mutant protein altered its structural stability and increased the binding energy with PPP1CB, indicating unstable protein-protein interactions. On WB, a low-molecular-weight band was observed, indicating that the protein was truncated. Co-IP indicated that the mutant protein no longer interacted with PPP1CB, indicating an effect on the structural stability of the myosin phase complex. CONCLUSION: The PPP1R12A c.2533 C > T variant may explain the neonatal seizures in the present case. The findings of this study expand the spectrum of PPP1R12A variants and highlight the potential significance of truncated proteins in the pathogenesis of GUBS.

Female↗

ELISA: a unified, multidimensional view of the protein domain universe.

ELISA (http://romi.bu.edu/elisa/) is a database that was designed for flexibility in defining interesting queries about protein domain evolution. We have defined and included both the inherent characteristics of the domains such as structure and function and comparisons of these characteristics between domains. Thus, the database is useful in defining structural and functional links between related protein domains and by extension sequences that encode them. In this database we introduce and employ a novel method of functional annotation and comparison. For each protein domain we create a probabilistic functional annotation tree using GO. We have designed an algorithm that accurately compares these trees and thus provides a measure of "functional distance" between two protein domains. Along with functional annotation, we have also included structural comparison between protein domains and best sequence comparisons to all known genomes. The latter enables researchers to dynamically do searches for domains sharing similar phylogenetic profiles. This combination of data and tools enables the researcher to design complex queries to carry out research in the areas of protein domain evolution, structure prediction and functional annotation of novel sequences.

Amino Acid Sequence↗

[Construction and analysis of a metagenomic library from Tengchong hot spring soil in Yunnan Province].

By a combination of freezing/thawing/ proteinase K-based method and SDS/high-salt/heating treatment, the mixed environmental genomic DNA was isolated directly from a hot spring soil in Tengchong, Yunnan, China. With this method, The DNA yield was up to 1 - 2 microg/g soil. After purification with the Wizard DNA clean up system (Promega, Madison, Wis), the mixed genomic DNA was partially digested with restriction enzyme Pst I. Digested DNA fragments of 3 - 8 kb were recovered from agrose gel and ligated to the pSK (+) vector. The ligation mixture was transformed into DH10B strain, resulting in the construction of a metagenomic library with about 2.5 x 10(4) clones. Restriction enzyme analysis revealed that the average insert is about 4.6 kb. Some novel sequences were identified via sequencing and gene annotation analysis of 30 clones randomly chosen from this library.

China↗

Development and application of a salmonid EST database and cDNA microarray: data mining and interspecific hybridization characteristics.

We report 80,388 ESTs from 23 Atlantic salmon (Salmo salar) cDNA libraries (61,819 ESTs), 6 rainbow trout (Oncorhynchus mykiss) cDNA libraries (14,544 ESTs), 2 chinook salmon (Oncorhynchus tshawytscha) cDNA libraries (1317 ESTs), 2 sockeye salmon (Oncorhynchus nerka) cDNA libraries (1243 ESTs), and 2 lake whitefish (Coregonus clupeaformis) cDNA libraries (1465 ESTs). The majority of these are 3' sequences, allowing discrimination between paralogs arising from a recent genome duplication in the salmonid lineage. Sequence assembly reveals 28,710 different S. salar, 8981 O. mykiss, 1085 O. tshawytscha, 520 O. nerka, and 1176 C. clupeaformis putative transcripts. We annotate the submitted portion of our EST database by molecular function. Higher- and lower-molecular-weight fractions of libraries are shown to contain distinct gene sets, and higher rates of gene discovery are associated with higher-molecular weight libraries. Pyloric caecum library group annotations indicate this organ may function in redox control and as a barrier against systemic uptake of xenobiotics. A microarray is described, containing 7356 salmonid elements representing 3557 different cDNAs. Analyses of cross-species hybridizations to this cDNA microarray indicate that this resource may be used for studies involving all salmonids.

Animals↗

Cloning, sequencing, and biochemical characterization of the nostocyclopeptide biosynthetic gene cluster: molecular basis for imine macrocyclization.

Nostocyclopeptides A1 and A2 are novel cyclic heptapeptides produced by the terrestrial cyanobacterium Nostoc sp. ATCC53789 that possess a unique imino linkage in the macrocyclic ring. Herein we report the cloning, sequencing, annotation, and biochemical analysis of the 33-kb nostocyclopeptide (ncp) biosynthetic gene cluster, which includes seven open reading frames predicted to be involved in the biosynthesis and transport of these natural products. The genetic architecture and domain organization of the ncpA-B nonribosomal peptide synthetase (NRPS) is co-linear in arrangement with respect to the putative order of the biosynthetic assembly of the cyclic peptide. A reductase domain identified at the C-terminal end of the NRPS NcpB is predicted to catalyze an NAD(P)H-mediated hydride transfer to the heptapeptidyl-S-enzyme intermediate NH(2)-Tyr-Gly-DGln-Ile-Ser-mPro-Leu/Phe-S-NRPS to yield a linear heptapeptide aldehyde that is subsequently captured intramolecularly with the amino group of the N-terminal amino acid residue tyrosine to form a stable imine bond. While a few C-terminal reductases associated with NRPSs have been identified, the ncp reductase is the first to mediate imine macrocyclization involving peptide N- and C-termini. Biochemical analysis of the NcpA1 and NcpB1 adenylation domains coupled with the recent characterization of the (2S,4S)-5-hydroxyleucine dehydrogenase NcpD, which is involved in the biosynthesis of the nonproteinogenic amino acid residue L-4-methylproline from L-leucine, support the involvement of this cluster in nostocyclopeptide biosynthesis.

Cloning, Molecular↗

Phylomat: an automated protein motif analysis tool for phylogenomics.

Recent progress in genomics, proteomics, and bioinformatics enables unprecedented opportunities to examine the evolutionary history of molecular, cellular, and developmental pathways through phylogenomics. Accordingly, we have developed a motif analysis tool for phylogenomics (Phylomat, http://alg.ncsa.uiuc.edu/pmat) that scans predicted proteome sets for proteins containing highly conserved amino acid motifs or domains for in silico analysis of the evolutionary history of these motifs/domains. Phylomat enables the user to download results as full protein or extracted motif/domain sequences from each protein. Tables containing the percent distribution of a motif/domain in organisms normalized to proteome size are displayed. Phylomat can also align the set of full protein or extracted motif/domain sequences and predict a neighbor-joining tree from relative sequence similarity. Together, Phylomat serves as a user-friendly data-mining tool for the phylogenomic analysis of conserved sequence motifs/domains in annotated proteomes from the three domains of life.

Algorithms↗

Predicting function: from genes to genomes and back.

Predicting function from sequence using computational tools is a highly complicated procedure that is generally done for each gene individually. This review focuses on the added value that is provided by completely sequenced genomes in function prediction. Various levels of sequence annotation and function prediction are discussed, ranging from genomic sequence to that of complex cellular processes. Protein function is currently best described in the context of molecular interactions. In the near future it will be possible to predict protein function in the context of higher order processes such as the regulation of gene expression, metabolic pathways and signalling cascades. The analysis of such higher levels of function description uses, besides the information from completely sequenced genomes, also the additional information from proteomics and expression data. The final goal will be to elucidate the mapping between genotype and phenotype.

Bacterial Proteins↗

Walking through protein sequence space.

Following the original idea of Maynard Smith on evolution of the protein sequence space, a novel tool is developed that allows the "space walk", from one sequence to its likely evolutionary relative and further on. At a given threshold of identity between consecutive steps, the walks of many steps are possible. The sequences at the ends of the walks may substantially differ from one another. In a sequence space of randomized (shuffled) sequences the walks are very short. The approach opens new perspectives for protein evolutionary studies and sequence annotation.

ATP-Binding Cassette Transporters↗

ProteoParc: A Reference Protein Database Builder for Ancient and Nonmodel Organisms.

Over the past few years, the increasing interest in analyzing the proteome of extinct and nonmodel organisms has generated a new field of research expanding the scope of proteomics. The lack of curated databases and/or molecular data from these organisms forces researchers to manually search in different public repositories for related protein sequences, either for MS/MS peptide identification or ZooMS marker annotation. This can lead to format incongruences and hinder reproducibility between studies. To address this issue, we introduce ProteoParc, a user-friendly software that builds reference databases by systematically downloading and processing protein sequences from the most widely used public repositories. The pipeline's output is a nonredundant protein database, formatted in a way to be interpreted by typical peptide identification software. Moreover, the user can adjust the database dimension and composition by applying different criteria to include only a certain number of genes or species. Thus, ProteoParc is an easy and fast, custom-made bioinformatic tool useful for future paleoproteomics analysis in ancient samples related to understudied organisms.

Databases, Protein↗

Respiratory hydrogen use by Salmonella enterica serovar Typhimurium is essential for virulence.

Based on available annotated gene sequence information, the enteric pathogen salmonella, like other enteric bacteria, contains three putative membrane-associated H2-using hydrogenase enzymes. These enzymes split molecular H2, releasing low-potential electrons that are used to reduce quinone or heme-containing components of the respiratory chain. Here we show that each of the three distinct membrane-associated hydrogenases of Salmonella enterica serovar Typhimurium is coupled to a respiratory pathway that uses oxygen as the terminal electron acceptor. Cells grown in a blood-based medium expressed four times the amount of hydrogenase (H2 oxidation) activity that cells grown on Luria Bertani medium did. Cells suspended in phosphate-buffered saline consumed 2 mol of H2 per mol of O2 used in the H2-O2 respiratory pathway, and the activity was inhibited by the respiration inhibitor cyanide. Molecular hydrogen levels averaging over 40 microM were measured in organs (i.e., livers and spleens) of live mice, and levels within the intestinal tract (the presumed origin of the gas) were four times greater than this. The half-saturation affinity of S. enterica serovar Typhimurium for H2 is only 2.1 microM, so it is expected that H2-utilizing hydrogenase enzymes are saturated with the reducing substrate in vivo. All three hydrogenase enzymes contribute to the virulence of the bacterium in a typhoid fever-mouse model, based on results from strains with mutations in each of the three hydrogenase genes. The introduced mutations are nonpolar, and growth of the mutant strains was like that of the parent strain. The combined removal of all three hydrogenases resulted in a strain that is avirulent and (in contrast to the parent strain) one that is unable to invade liver or spleen tissue. The introduction of one of the hydrogenase genes into the triple mutant strain on a low-copy-number plasmid resulted in a strain that was able to both oxidize H2 and cause morbidity in mice within 11 days of inoculation; therefore, the avirulent phenotype of the triple mutant is not due to an unknown spurious mutation. We conclude that H2 utilization in a respiratory fashion is required for energy production to permit salmonella growth and subsequent virulence during infection.

Animals↗

VKCDB: voltage-gated potassium channel database.

BACKGROUND: The family of voltage-gated potassium channels comprises a functionally diverse group of membrane proteins. They help maintain and regulate the potassium ion-based component of the membrane potential and are thus central to many critical physiological processes. VKCDB (Voltage-gated potassium [K] Channel DataBase) is a database of structural and functional data on these channels. It is designed as a resource for research on the molecular basis of voltage-gated potassium channel function. DESCRIPTION: Voltage-gated potassium channel sequences were identified by using BLASTP to search GENBANK and SWISSPROT. Annotations for all voltage-gated potassium channels were selectively parsed and integrated into VKCDB. Electrophysiological and pharmacological data for the channels were collected from published journal articles. Transmembrane domain predictions by TMHMM and PHD are included for each VKCDB entry. Multiple sequence alignments of conserved domains of channels of the four Kv families and the KCNQ family are also included. Currently VKCDB contains 346 channel entries. It can be browsed and searched using a set of functionally relevant categories. Protein sequences can also be searched using a local BLAST engine. CONCLUSIONS: VKCDB is a resource for comparative studies of voltage-gated potassium channels. The methods used to construct VKCDB are general; they can be used to create specialized databases for other protein families. VKCDB is accessible at http://vkcdb.biology.ualberta.ca.

Animals↗

A Molecularly Anchored Spatial Transcriptomic Framework for Precise CA1-Subiculum Parcellation and Region-Resolved Analysis in Alzheimer's Disease.

BACKGROUND: The precise molecular delineation of the interface between the Subiculum (Sub) and cornu ammonis 1 (CA1) is a challenge in hippocampal research, as conventional cytoarchitectural boundaries are often ambiguous and limit reproducible regional annotation. Here, we developed a molecularly anchored spatial transcriptomic framework to define CA1-Sub regional identities using high-definition spatial transcriptomics (Stereo-seq) and single-nucleus RNA sequencing (snRNA-seq) references. FINDINGS: Using a human hippocampal Stereo-seq dataset from 12 donors, we established a data-driven parcellation framework that defines reproducible molecular features distinguishing CA1 and Sub while capturing the transition between these regions. FN1 was identified as a Sub-enriched marker in a subset of EX_Sub and, together with ETV1 and additional regional markers, enabled molecular assignment of CA1 and Sub identities across datasets. The Sub association of FN1 and ETV1 was further supported by human 10X Genomics spatial transcriptomics, mouse in situ hybridization data, and a mouse spatial transcriptomic dataset. Applying this framework to Alzheimer's disease (AD) tissues revealed region-specific transcriptional alterations across CA1 and Sub, including enrichment of mitochondrial energy metabolism-related transcripts in the Sub, suggesting exploratory transcriptional associations of altered metabolic function. CONCLUSIONS: This study provides a molecularly anchored framework for human CA1-Sub parcellation that complements conventional annotation. By defining regional molecular states while preserving the biological continuum across CA1-Sub interface, this approach enables more consistent regional analysis of human hippocampus tissue across donors, datasets, and disease conditions.

Journal Article↗

Microarray gene expression profiles in dilated and hypertrophic cardiomyopathic end-stage heart failure.

Despite similar clinical endpoints, heart failure resulting from dilated cardiomyopathy (DCM) or hypertrophic cardiomyopathy (HCM) appears to develop through different remodeling and molecular pathways. Current understanding of heart failure has been facilitated by microarray technology. We constructed an in-house spotted cDNA microarray using 10,272 unique clones from various cardiovascular cDNA libraries sequenced and annotated in our laboratory. RNA samples were obtained from left ventricular tissues of precardiac transplantation DCM and HCM patients and were hybridized against normal adult heart reference RNA. After filtering, differentially expressed genes were determined using novel analyzing software. We demonstrated that normalization for cDNA microarray data is slide-dependent and nonlinear. The feasibility of this model was validated by quantitative real-time reverse transcription-PCR, and the accuracy rate depended on the fold change and statistical significance level. Our results showed that 192 genes were highly expressed in both DCM and HCM (e.g., atrial natriuretic peptide, CD59, decorin, elongation factor 2, and heat shock protein 90), and 51 genes were downregulated in both conditions (e.g., elastin, sarcoplasmic/endoplasmic reticulum Ca2+-ATPase). We also identified several genes differentially expressed between DCM and HCM (e.g., alphaB-crystallin, antagonizer of myc transcriptional activity, beta-dystrobrevin, calsequestrin, lipocortin, and lumican). Microarray technology provides us with a genomic approach to explore the genetic markers and molecular mechanisms leading to heart failure.

Adult↗

"Plus-C" odorant-binding protein genes in two Drosophila species and the malaria mosquito Anopheles gambiae.

Olfaction plays a crucial role in many aspects of insect behaviour, including host selection by agricultural pests and vectors of human disease. Insect odorant-binding proteins (OBPs) are thought to function as the first step in molecular recognition and the transport of semiochemicals. The whole genome sequence of the fruit fly Drosophila melanogaster has been completed and a large number of genes have been annotated as OBPs, based on the presence of six conserved cysteine residues and a conserved spacing between the cysteines. These proteins can be divided into three distinct subgroups; those with only one six-cysteine motif, those with two such motifs and those with one motif, three extra conserved cysteines and a conserved proline immediately after the sixth cysteine. This study concentrates on the last two subgroups, referred to as 'dimer' OBPs and 'Plus-C' OBPs, respectively. We determined the tissue-specific transcript levels of all of these OBP genes of D. melanogaster using semiquantitative RT-PCR. The results showed that the expression patterns can vary within a subgroup of genes and that this technique is valuable for assessing which of the putative OBP genes are likely to be involved in Drosophila olfaction. The publicly available genomes of another fruit fly Drosophila pseudoobscura, the malaria mosquito Anopheles gambiae and the yellow fever mosquito Aedes aegypti were searched by Blast against each Plus-C OBP and dimer OBP of D. melanogaster. Related genes were found in all of the other species and the relationships of these with the D. melanogaster genes and their possible biological functions are discussed.

Amino Acid Sequence↗

The nematode Caenorhabditis elegans as a model to study the roles of proteoglycans.

The nematode Caenorhabditis elegans is a powerful animal model for exploring the genetic basis of metazoan development. Recent genetic and biochemical studies have revealed that the molecular machinery of glycosaminoglycan (GAG) biosynthesis and modification is highly conserved between C. elegans and mammals. In addition, genetic studies have implicated GAGs in vulval morphogenesis and zygotic cytokinesis. The extensive knowledge of C. elegans biology, including its elucidated cell lineage, together with the completed and well annotated DNA sequence and availability of reverse genetic tools, provide a platform for studying the functions of proteoglycans and their GAG modification.

Animals↗

Molecular genetic analyses of potential beta-galactosidase genes in Xanthomonas campestris.

Xanthomonas campestris pv. campestris, which displays no significant beta-1,4-D-galactopyranosidase activity, has three annotated beta-galactosidase genes in the sequenced genome, designated galA, galB and galC herein. GalA and GalB are similar to glycosyl hydrolase (GH) family 2 enzymes, including Escherichia coli LacZ. galA and galB cannot express detectable activity even after being cloned in-frame and driven by the vector's promoter. GalC is a GH35 enzyme homologous to the Xanthomonas axonopodis pv. manihotis Bga. The latter cleaves beta,1-3-linked galactose 1,000 times faster than beta,1-4-linked galactose and is not responsible for lactose utilization. In X. campestris pv. campestris cells, GalC is readily detectable by Western blotting, and the levels can be increased by cloning the gene under the control of the vector's promoter. Results of insertional mutation, transcriptional fusion assay and Western blotting indicated that galC, clustered with several GH genes, is cotranscribed with the upstream gene(s) and is expressed constitutively. Xc17L is a previously isolated mutant with elevated beta-galactosidase activity and a greatly improved ability to grow on lactose. Results of DNA sequencing of Xc17L galA, galB and galC, enzyme assays of galA, galB and galC mutants derived from Xc17L, and Western blotting of GalC in Xc17L indicated that the three beta-galactosidase genes do not encode the elevated beta-galactosidase activity in Xc17L. The presence of a fourth beta-galactosidase gene is proposed.

Amino Acid Motifs↗

FISH analysis of Drosophila melanogaster heterochromatin using BACs and P elements.

The heterochromatin of chromosomes 2 and 3 of Drosophila melanogaster contains about 30 essential genes defined by genetic analysis. In the last decade only a few of these genes have been molecularly characterized and found to correspond to protein-coding genes involved in important cellular functions. Moreover, several predicted genes have been identified by annotation of genomic sequence that are associated with polytene chromosome divisions 40, 41 and 80 but their locations on the cytogenetic map of the heterochromatin are still uncertain. To expand our current knowledge of the genetic functions located in heterochromatin, we have performed fluorescence in situ hybridization (FISH) mapping to mitotic chromosomes of nine bacterial artificial chromosomes (BACs) carrying several predicted genes and of 13 P element insertions assigned to the proximal regions of 2R and 3L. We found that 22 predicted genes map to the h46 region of 2R and eight map to the h47 regions of 3L. This amounts to at least 30 predicted genes located in these heterochromatic regions, whereas previous studies detected only seven vital genes. Finally, another 58 genes localize either in the euchromatin-heterochromatin transition regions or in the proximal euchromatin of 2R and 3L.

Animals↗