Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Molecular Sequence Annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Primer on medical genomics. Part IV: Expression proteomics.

Proteomics, simply defined, is the study of proteomes. More completely, proteomics is defined as the study of all proteins, including their relative abundance, distribution, posttranslational modifications, functions, and interactions with other macromolecules, in a given cell or organism within a given environment and at a specific stage in the cell cycle. Proteins carry out the biological functions encoded by genes; hence, once the initial stage of genome sequencing and gene discovery is completed, a study of the proteome must be undertaken to address fundamental biological questions. The 3 broad areas are expression proteomics, which catalogues the relative abundance of proteins; cell-mapping or cellular proteomics, which delineates functional protein-protein interactions and organelle-specific protein distribution; and structural proteomics, which characterizes the 3-dimensional structure of proteins. With these approaches, proteins are studied on a global scale using a synergistic combination of powerful, high-throughput technologies, including 2-dimensional polyacrylamide gel electrophoresis, mass spectrometry, multidimensional liquid chromatography, and bioinformatics. Mass spectrometry, which provides highly accurate molecular mass measurements, has emerged as the analytical technology of choice for protein identification, characterization, and sequencing. This task has been made considerably easier with the availability of complete, nonredundant, and annotated genome sequence databases for many organisms. This article reviews the area of expression proteomics.

Biotechnology↗

Genetic landscape of pediatric seizures in Southeast China: identification of a novel GLI3 frameshift variant through whole-exome sequencing.

BACKGROUND: Pediatric seizure disorders are clinically and genetically heterogeneous. Whole-exome sequencing has improved the detection of rare genetic variants in childhood epilepsy; however, data from pediatric populations in Southeast China remain limited. This study aimed to characterize the genetic landscape of pediatric seizure disorders in Southeast China and to evaluate the clinical diagnostic yield of whole-exome sequencing. MATERIALS AND METHODS: This retrospective observational study included 21 pediatric patients with seizure disorders who were recruited at the Fifth Hospital of Xiamen, Fujian, China, between January 2021 and June 2024. Clinical data were extracted from medical records. Whole-exome sequencing was performed on DNA extracted from peripheral blood. Sequence variants were annotated, filtered, and classified according to the guidelines of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology. Copy-number variants were evaluated using exome-based algorithms. Descriptive statistics were used because of the limited sample size. RESULTS: WES identified three clinically relevant, likely pathogenic findings in 3 of 21 patients, corresponding to a provisional diagnostic yield of 14.3%. The remaining 62 of 65 variants were of uncertain significance (VUS). The three retained variants included a GLI3 frameshift variant (exon 2: c.90_91insCAGATGTGAGC; p.Glu31Glnfs*3) and two copy-number variants (16p13.12-16p13.11 duplication and Xp22.31 deletion) with established clinical significance. Functional analysis of all 65 variants revealed that ion channel genes and neurodevelopmental genes were the most frequently affected categories. CONCLUSION: Whole-exome sequencing identified clinically relevant genetic findings in a subset of Southeast Chinese children with seizure disorders. The novel GLI3 frameshift variant may suggest an expansion of the GLI3-associated phenotypic spectrum, but further segregation, functional validation, and larger cohort studies are needed. The high proportion of variants of uncertain significance highlights the ongoing challenges of genetic interpretation in pediatric seizure disorders.

GLI3 frameshift variant↗

Cataloging transcription factor and major signaling molecule genes for functional genomic studies in Ciona intestinalis.

The ascidian Ciona intestinalis provides an excellent experimental system for functional genomic studies because (1) its genome has been sequenced, (2) the transcription factor genes and genes for major signal transduction molecules have been extensively screened and annotated on a genome-wide scale using the molecular phylogenetical method, and (3) their embryonic expression profiles have been almost completely determined. However, the entire genetic structure, including the 5' and 3' untranslated regions and the protein-coding regions, of most gene models used in these prior studies is not always supported by cDNA evidence, and thus, these gene models are potentially imprecise. To facilitate functional genomic studies based on precise gene structures, our present study determined 406 cDNA sequences for 357 transcription factor genes and 112 cDNA sequences for 107 signal transduction molecule genes, greatly improving the previous gene models and revealing transcript variants for 44 genes. Considering these data alongside those of previously characterized genes deposited in the DNA Data Bank of Japan/European Molecular Biology Laboratory/GENBANK databases, 95.6% of the catalogued transcription factor genes (373/390) and 98.3% of the catalogued signal transduction molecule genes (117/119) have now been verified by cDNA sequences. Thus, the present study greatly improves the resources available for functional genomic studies in C. intestinalis.

Animals↗

Survey of transcripts in the adult Drosophila brain.

BACKGROUND: Classic methods of identifying genes involved in neural function include the laborious process of behavioral screening of mutagenized flies and then rescreening candidate lines for pleiotropic effects due to developmental defects. To accelerate the molecular analysis of brain function in Drosophila we constructed a cDNA library exclusively from adult brains. Our goal was to begin to develop a catalog of transcripts expressed in the brain. These transcripts are expected to contain a higher proportion of clones that are involved in neuronal function. RESULTS: The library contains approximately 6.75 million independent clones. From our initial characterization of 271 randomly chosen clones, we expect that approximately 11% of the clones in this library will identify transcribed sequences not found in expressed sequence tag databases. Furthermore, 15% of these 271 clones are not among the 13,601 predicted Drosophila genes. CONCLUSIONS: Our analysis of this unique Drosophila brain library suggests that the number of genes may be underestimated in this organism. This work complements the Drosophila genome project by providing information that facilitates more complete annotation of the genomic sequence. This library should be a useful resource that will help in determining how basic brain functions operate at the molecular level.

Aging↗

A rigorous method for multigenic families' functional annotation: the peptidyl arginine deiminase (PADs) proteins family example.

BACKGROUND: large scale and reliable proteins' functional annotation is a major challenge in modern biology. Phylogenetic analyses have been shown to be important for such tasks. However, up to now, phylogenetic annotation did not take into account expression data (i.e. ESTs, Microarrays, SAGE, ...). Therefore, integrating such data, like ESTs in phylogenetic annotation could be a major advance in post genomic analyses. We developed an approach enabling the combination of expression data and phylogenetic analysis. To illustrate our method, we used an example protein family, the peptidyl arginine deiminases (PADs), probably implied in Rheumatoid Arthritis. RESULTS: the analysis was performed as follows: we built a phylogeny of PAD proteins from the NCBI's NR protein database. We completed the phylogenetic reconstruction of PADs using an enlarged sequence database containing translations of ESTs contigs. We then extracted all corresponding expression data contained in EST database This analysis allowed us 1/To extend the spectrum of homologs-containing species and to improve the reconstruction of genes' evolutionary history. 2/To deduce an accurate gene expression pattern for each member of this protein family. 3/To show a correlation between paralogous sequences' evolution rate and pattern of tissular expression. CONCLUSION: coupling phylogenetic reconstruction and expression data is a promising way of analysis that could be applied to all multigenic families to investigate the relationship between molecular and transcriptional evolution and to improve functional annotation.

Animals↗

Compilation of all genes encoding two-component phosphotransfer signal transducers in the genome of Escherichia coli.

Bacteria have devised sophisticated His-Asp phosphorelay signaling systems for eliciting a variety of adaptive responses to their environment, which are generally referred to as the "two-component regulatory system." The widespread occurrence of the His-Asp phosphorelay signaling in both prokaryotes and eukaryotes implies that it is a powerful device for a wide variety of adaptive responses of cells to their environment. The two-component signal transducers contain one or more of three common and characteristic phosphotransfer signaling domains, named the "transmitter, receiver, and histidine-containing phosphotransfer (HPt) domains." The recently determined entire genomic sequence of Escherichia coli allowed us to compile systematically a complete list of genes encoding such two-component signal transduction proteins. The results of such an effort, made in this study, revealed that at least 62 open reading frames (ORFs) were identified as putative members of the two-component signal transducers in this single species. Among them, 32 were identified as response regulator and 23 were identified as orthodox sensory kinases. In addition, E. coli has five hybrid sensory kinases. The precise location of each ORF was mapped on a physical map of the entire E. coli genome. All of these ORFs were then compiled and annotated extensively.

Amino Acid Sequence↗

Genomic organization and molecular characterization of Clostridium difficile bacteriophage PhiCD119.

In this study, we have isolated a temperate phage (PhiCD119) from a pathogenic Clostridium difficile strain and sequenced and annotated its genome. This virus has an icosahedral capsid and a contractile tail covered by a sheath and contains a double-stranded DNA genome. It belongs to the Myoviridae family of the tailed phages and the order Caudovirales. The genome was circularly permuted, with no physical ends detected by sequencing or restriction enzyme digestion analysis, and lacked a cos site. The DNA sequence of this phage consists of 53,325 bp, which carries 79 putative open reading frames (ORFs). A function could be assigned to 23 putative gene products, based upon bioinformatic analyses. The PhiCD119 genome is organized in a modular format, which includes modules for lysogeny, DNA replication, DNA packaging, structural proteins, and host cell lysis. The PhiCD119 attachment site attP lies in a noncoding region close to the putative integrase (int) gene. We have identified the phage integration site on the C. difficile chromosome (attB) located in a noncoding region just upstream of gene gltP, which encodes a carrier protein for glutamate and aspartate. This genetic analysis represents the first complete DNA sequence and annotation of a C. difficile phage.

Bacteriophages↗

PA-GOSUB: a searchable database of model organism protein sequences with their predicted Gene Ontology molecular function and subcellular localization.

PA-GOSUB (Proteome Analyst: Gene Ontology Molecular Function and Subcellular Localization) is a publicly available, web-based, searchable and downloadable database that contains the sequences, predicted GO molecular functions and predicted subcellular localizations of more than 107,000 proteins from 10 model organisms (and growing), covering the major kingdoms and phyla for which annotated proteomes exist (http://www.cs.ualberta.ca/~bioinfo/PA/GOSUB). The PA-GOSUB database effectively expands the coverage of subcellular localization and GO function annotations by a significant factor (already over five for subcellular localization, compared with Swiss-Prot v42.7), and more model organisms are being added to PA-GOSUB as their sequenced proteomes become available. PA-GOSUB can be used in three main ways. First, a researcher can browse the pre-computed PA-GOSUB annotations on a per-organism and per-protein basis using annotation-based and text-based filters. Second, a user can perform BLAST searches against the PA-GOSUB database and use the annotations from the homologs as simple predictors for the new sequences. Third, the whole of PA-GOSUB can be downloaded in either FASTA or comma-separated values (CSV) formats.

Amino Acid Sequence↗

Biosequence exegesis.

Annotation of large-scale gene sequence data will benefit from comprehensive and consistent application of well-documented, standard analysis methods and from progressive and vigilant efforts to ensure quality and utility and to keep the annotation up to date. However, it is imperative to learn how to apply information derived from functional genomics and proteomics technologies to conceptualize and explain the behaviors of biological systems. Quantitative and dynamical models of systems behaviors will supersede the limited and static forms of single-gene annotation that are now the norm. Molecular biological epistemology will increasingly encompass both teleological and causal explanations.

Animals↗

Multiomic study of cutaneous T-cell lymphoma reveals single-cell clonal evolution in progression and therapy resistance.

Cutaneous T-cell lymphoma (CTCL) remains a challenging disease due to its significant heterogeneity, therapy resistance, and relentless progression. Multiomics technologies offer the potential to provide uniquely precise views of disease progression and response to therapy. Here, we present a comprehensive multiomics view of CTCL clonal evolution, incorporating exome, whole-genome, epigenome, bulk, single-cell T-cell receptor, and single-cell RNA sequencing of 99 clinically annotated serial skin, peripheral blood, and lymph node samples from 34 patients with CTCL. We leveraged this extensive data set to define the molecular underpinnings of CTCL progression in individual patients at single-cell resolution with the goal of identifying clinically useful biomarkers and therapeutic targets. Our studies identified recurrent progression-associated clonal genomic alterations; we highlight mutation of CCR4, phosphoinositide 3-kinase inhibitor signaling, and programmed cell death protein 1 (PD-1) checkpoint pathways as evasion tactics deployed by malignant T cells. We identified a gain-of-function mutation in STAT3 (D661Y) and demonstrated, using cleavage under targets and release using nuclease (CUT&RUN) and RNA sequencing, that it enhances binding to and transcription of genes in Rho GTPase pathways. With our previous work implicating this pathway in histone deacetylase inhibitor-resistant CTCL, these data provide further support for a previously unrecognized role for Rho GTPase pathway dysregulation in CTCL progression. Recurrent progression-associated mutations were common in the epigenetic modifier EZH2, suggesting that EZH2 inhibition may benefit patients with CTCL. Our findings support an approach in which genomic analysis is widely used for improved disease monitoring, biomarker-informed clinical trial design, and genome-guided therapeutic decision-making. Moreover, these molecular changes present new opportunities for therapeutic targeting in this challenging and incurable cancer.

Multiomics↗

iProLINK: an integrated protein resource for literature mining.

The exponential growth of large-scale molecular sequence data and of the PubMed scientific literature has prompted active research in biological literature mining and information extraction to facilitate genome/proteome annotation and improve the quality of biological databases. Motivated by the promise of text mining methodologies, but at the same time, the lack of adequate curated data for training and benchmarking, the Protein Information Resource (PIR) has developed a resource for protein literature mining--iProLINK (integrated Protein Literature INformation and Knowledge). As PIR focuses its effort on the curation of the UniProt protein sequence database, the goal of iProLINK is to provide curated data sources that can be utilized for text mining research in the areas of bibliography mapping, annotation extraction, protein named entity recognition, and protein ontology development. The data sources for bibliography mapping and annotation extraction include mapped citations (PubMed ID to protein entry and feature line mapping) and annotation-tagged literature corpora. The latter includes several hundred abstracts and full-text articles tagged with experimentally validated post-translational modifications (PTMs) annotated in the PIR protein sequence database. The data sources for entity recognition and ontology development include a protein name dictionary, word token dictionaries, protein name-tagged literature corpora along with tagging guidelines, as well as a protein ontology based on PIRSF protein family names. iProLINK is freely accessible at http://pir.georgetown.edu/iprolink, with hypertext links for all downloadable files.

Computational Biology↗

FusionTarget: Computational framework for drug repurposing against modeled fusion protein structures from genomic breakpoints.

Many fusion genes have been recognized as biomarkers and therapeutic targets. However, the lack of knowledge on protein structures and targeting approaches made it challenging to develop effective targeting therapeutics. To fill this, we developed a computational pipeline, FusionTarget, which annotates the genomic DNA breakage to RNA and protein sequences, predicts the 3D structures of fusion proteins, and performs comparative virtual screening, comparative molecular dynamics simulation, and quantitative analyses to identify the fusion protein-selective small molecules by selecting drugs with consistent high-fold binding affinity between fusion and wild-type proteins in multiple isoforms. We applied our pipeline to EWSR1::FLI1 in Ewing sarcoma and KMT2A::AFF1 in infant acute lymphoblastic leukemia. Further cell assay experiments confirmed that cells expressing individual fusion genes were more sensitive to the suggested drugs, and the key downstream genes were affected by our drugs. FusionTarget provides a unique foundation for developing therapeutics targeting fusion proteins.

applied computing in medical science↗

Genome-wide ENU mutagenesis to reveal immune regulators.

A complete list of molecular components for immune system function is now available with the completion of the human and mouse genome sequences. However, identification and functional annotation of genes involved in immunological processes require a discovery methodology that can efficiently and broadly analyze the complex interplay of these components in vivo. Our recent experience indicates that genome-wide chemical mutagenesis in the mouse is an extremely powerful methodology for the identification of genes required for complex immunological processes.

Animals↗

PlantGDB, plant genome database and analysis tools.

PlantGDB (http://www.plantgdb.org/) is a database of molecular sequence data for all plant species with significant sequencing efforts. The database organizes EST sequences into contigs that represent tentative unique genes. Contigs are annotated and, whenever possible, linked to their respective genomic DNA. Genome sequence fragments are assembled similarly. The goal of the PlantGDB web site is to establish the basis for identifying sets of genes common to all plants or specific to particular species by integrating a number of bioinformatics tools that facilitate gene prediction and cross- species comparisons. For species with large-scale genome sequencing efforts, PlantGDB provides genome browsing capabilities that integrate all available EST and cDNA evidence for current gene models (for Arabidopsis thaliana, see the AtGDB site at http://www.plantgdb.org/AtGDB/).

Computational Biology↗

Genomic exploration and in silico prioritization of putative COX-2-targeting metabolites from Streptomyces sp. VITGV156 (MCC 4965).

INTRODUCTION: Streptomyces species represent an important source of bioactive natural products, yet systematic genome-guided prioritization of metabolites targeting cyclooxygenase-2 (COX-2/PTGS2) remains limited. This study aimed to investigate the biosynthetic potential of Streptomyces sp. VITGV156 (MCC 4965) using an integrated genome mining and computational drug discovery pipeline. METHODS: Whole-genome sequencing, functional annotation, antiSMASH v7.0.1-based biosynthetic gene cluster (BGC) prediction, LC-MS/MS metabolomic profiling, SwissADME analysis, target prediction, disease association mapping, molecular docking against PTGS2 (PDB: 5IKR), and PASS bioactivity prediction were performed to prioritize putative bioactive metabolites. RESULTS: Genome analysis identified 29 predicted biosynthetic gene clusters, including clusters associated with geosmin, ectoine, albaflavenone, hopene, coelichelin, and SapB, together with several cryptic clusters exhibiting low similarity to known pathways. LC-MS/MS metabolomic profiling provided experimental support for active secondary metabolite production under the cultivation conditions employed. Computational prioritization identified PTGS2 (COX-2) as a biologically relevant target. Molecular docking demonstrated favorable binding affinities and interaction profiles for several predicted metabolites within the PTGS2 catalytic pocket. PASS analysis further suggested potential anticancer-related biological activities that require experimental validation. DISCUSSION: These findings demonstrate the utility of integrating genome mining, metabolomic profiling, and computational drug discovery for prioritizing natural-product candidates. Streptomyces sp. VITGV156 (MCC 4965) represents a promising source of biosynthetic diversity and provides a genome-guided framework for identifying putative COX-2-targeting natural products for future experimental validation rather than confirming metabolite production or biological activity.

COX-2 (PTGS2)↗

Cyanidioschyzon merolae genome. A tool for facilitating comparable studies on organelle biogenesis in photosynthetic eukaryotes.

The ultrasmall unicellular red alga Cyanidioschyzon merolae lives in the extreme environment of acidic hot springs and is thought to retain primitive features of cellular and genome organization. We determined the 16.5-Mb nuclear genome sequence of C. merolae 10D as the first complete algal genome. BLASTs and annotation results showed that C. merolae has a mixed gene repertoire of plants and animals, also implying a relationship with prokaryotes, although its photosynthetic components were comparable to other phototrophs. The unicellular green alga Chlamydomonas reinhardtii has been used as a model system for molecular biology research on, for example, photosynthesis, motility, and sexual reproduction. Though both algae are unicellular, the genome size, number of organelles, and surface structures are remarkably different. Here, we report the characteristics of double membrane- and single membrane-bound organelles and their related genes in C. merolae and conduct comparative analyses of predicted protein sequences encoded by the genomes of C. merolae and C. reinhardtii. We examine the predicted proteins of both algae by reciprocal BLASTP analysis, KOG assignment, and gene annotation. The results suggest that most core biological functions are carried out by orthologous proteins that occur in comparable numbers. Although the fundamental gene organizations resembled each other, the genes for organization of chromatin, cytoskeletal components, and flagellar movement remarkably increased in C. reinhardtii. Molecular phylogenetic analyses suggested that the tubulin is close to plant tubulin rather than that of animals and fungi. These results reflect the increase in genome size, the acquisition of complicated cellular structures, and kinematic devices in C. reinhardtii.

Algal Proteins↗

An atlas of differential gene expression during early Xenopus embryogenesis.

We have carried out a large-scale, semi-automated whole-mount in situ hybridization screen of 8369 cDNA clones in Xenopus laevis embryos. We confirm that differential gene expression is prevalent during embryogenesis since 24% of the clones are expressed non-ubiquitously and 8% are organ or cell type specific marker genes. Sequence analysis and clustering yielded 723 unique genes displaying a differential expression pattern. Of these, 18% were already described in Xenopus, 47% have homologs and 35% are lacking significant sequence similarity in databases. Many of them encode known developmental regulators. We classified 363 of the 723 genes for which a Gene Ontology annotation for molecular function could be attributed and found 'DNA binding' and 'enzyme' the most represented terms. The most common protein domains encoded in these embryonic, differentially expressed genes are the homeobox and RNA Recognition Motif (RRM). Fifty-nine putative orthologs of human disease genes, and 254 organ or cell specific marker genes were identified. Markers were found for nasal placode and archenteron roof, organs for which a specific marker was previously unavailable. Markers were also found for novel subdomains of various other organs. The tissues for which most markers were found are muscle and epidermis. Expression of cell cycle regulators fell in two classes, containing proliferation-promoting and anti-proliferative genes, respectively. We identified 66 new members of the BMP4, chromatin, endoplasmic reticulum, and karyopherin synexpression groups, thus providing a first glimpse of their probable cellular roles. Cluster analysis of tissues to measure tissue relatedness yielded some unorthodox affinities besides expectable lineage relationships. In conclusion, this study represents an atlas of gene expression patterns, which reveals embryonic regionalization, provides novel marker genes, and makes predictions about the functional role of unknown genes.

Animals↗

The first archaeal ATP-dependent glucokinase, from the hyperthermophilic crenarchaeon Aeropyrum pernix, represents a monomeric, extremely thermophilic ROK glucokinase with broad hexose specificity.

An ATP-dependent glucokinase of the hyperthermophilic aerobic crenarchaeon Aeropyrum pernix was purified 230-fold to homogeneity. The enzyme is a monomeric protein with an apparent molecular mass of about 36 kDa. The apparent K(m) values for ATP and glucose (at 90 degrees C and pH 6.2) were 0.42 and 0.044 mM, respectively; the apparent V(max) was about 35 U/mg. The enzyme was specific for ATP as a phosphoryl donor, but showed a broad spectrum for phosphoryl acceptors: in addition to glucose, which showed the highest catalytic efficiency (k(cat)/K(m)), the enzyme also phosphorylates glucosamin, fructose, mannose, and 2-deoxyglucose. Divalent cations were required for maximal activity: Mg(2+), which was most effective, could partially be replaced with Co(2+), Mn(2+), and Ni(2+). The enzyme had a temperature optimum of at least 100 degrees C and showed significant thermostability up to 100 degrees C. The coding function of open reading frame (ORF) APE2091 (Y. Kawarabayasi, Y. Hino, H. Horikawa, S. Yamazaki, Y. Haikawa, K. Jin-no, M. Takahashi, M. Sekine, S. Baba, A. Ankai, H. Kosugi, A. Hosoyama, S. Fukui, Y. Nagai, K. Nishijima, H. Nakazawa, M. Takamiya, S. Masuda, T. Funahashi, T. Tanaka, Y. Kudoh, J. Yamazaki, N. Kushida, A. Oguchi, and H. Kikuchi, DNA Res. 6:83-101, 145-152, 1999), previously annotated as gene glk, coding for ATP-glucokinase of A. pernix, was proved by functional expression in Escherichia coli. The purified recombinant ATP-dependent glucokinase showed a 5-kDa higher molecular mass on sodium dodecyl sulfate-polyacrylamide gel electrophoresis, but almost identical kinetic and thermostability properties in comparison to the native enzyme purified from A. pernix. N-terminal amino acid sequence of the native enzyme revealed that the translation start codon is a GTG 171 bp downstream of the annotated start codon of ORF APE2091. The amino acid sequence deduced from the truncated ORF APE2091 revealed sequence similarity to members of the ROK family, which comprise bacterial sugar kinases and transcriptional repressors. This is the first report of the characterization of an ATP-dependent glucokinase from the domain of Archaea, which differs from its bacterial counterparts by its monomeric structure and its broad specificity for hexoses.

Adenosine Triphosphate↗