Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Molecular Sequence Annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Recurrent use of evolutionary importance for functional annotation of proteins based on local structural similarity.

The annotation of protein function has not kept pace with the exponential growth of raw sequence and structure data. An emerging solution to this problem is to identify 3D motifs or templates in protein structures that are necessary and sufficient determinants of function. Here, we demonstrate the recurrent use of evolutionary trace information to construct such 3D templates for enzymes, search for them in other structures, and distinguish true from spurious matches. Serine protease templates built from evolutionarily important residues distinguish between proteases and other proteins nearly as well as the classic Ser-His-Asp catalytic triad. In 53 enzymes spanning 33 distinct functions, an automated pipeline identifies functionally related proteins with an average positive predictive power of 62%, including correct matches to proteins with the same function but with low sequence identity (the average identity for some templates is only 17%). Although these template building, searching, and match classification strategies are not yet optimized, their sequential implementation demonstrates a functional annotation pipeline which does not require experimental information, but only local molecular mimicry among a small number of evolutionarily important residues.

Algorithms↗

Genetic linkage map and expression analysis of genes expressed in the lamellae of the edible basidiomycete Pleurotus ostreatus.

Pleurotus ostreatus is an industrially cultivated basidiomycete with nutritional and environmental applications. Its genome contains 35 Mbp organized in 11 chromosomes. There is currently available a genetic linkage map based predominantly on anonymous molecular markers complemented with the mapping of QTLs controlling growth rate and industrial productivity. To increase the saturation of the existing linkage maps, we have identified and mapped 82 genes expressed in the lamellae. Their manual annotation revealed that 34.1% of the lamellae-expressed and 71.5% of the lamellae-specific genes correspond to previously unknown sequences or to hypothetical proteins without a clearly established function. Furthermore, the expression pattern of some genes provides an experimental basis for studying gene regulation during the change from vegetative to reproductive growth. Finally, the identification of various differentially regulated genes involved in protein metabolism suggests the relevance of these processes in fruit body formation and maturation.

Chromosome Mapping↗

Interspecies conservation of gene order and intron-exon structure in a genomic locus of high gene density and complexity in Plasmodium.

A 13.6 kb contig of chromosome 5 of Plasmodium berghei, a rodent malaria parasite, has been sequenced and analysed for its coding potential. Assembly and comparison of this genomic locus with the orthologous locus on chromosome 10 of the human malaria Plasmodium falciparum revealed an unexpectedly high level of conservation of the gene organisation and complexity, only partially predicted by current gene-finder algorithms. Adjacent putative genes, transcribed from complementary strands, overlap in their untranslated regions, introns and exons, resulting in a tight clustering of both regulatory and coding sequences, which is unprecedented for genome organisation of PLASMODIUM: In total, six putative genes were identified, three of which are transcribed in gametocytes, the precursor cells of gametes. At least in the case of two multiple exon genes, alternative splicing and alternative transcription initiation sites contribute to a flexible use of the dense information content of this locus. The data of the small sample presented here indicate the value of a comparative approach for Plasmodium to elucidate structure, organisation and gene content of complex genomic loci and emphasise the need to integrate biological data of all Plasmodium species into the P.falciparum genome database and associated projects such as PlasmodB to further improve their annotation.

Alternative Splicing↗

The functional genomic distribution of protein divergence in two animal phyla: coevolution, genomic conflict, and constraint.

We compare the functional spectrum of protein evolution in two separate animal lineages with respect to two hypotheses: (1) rates of divergence are distributed similarly among functional classes within both lineages, indicating that selective pressure on the proteome is largely independent of organismic-level biological requirements; and (2) rates of divergence are distributed differently among functional classes within each lineage, indicating species-specific selective regimes impact genome-wide substitutional patterns. Integrating comparative genome sequence with data from tissue-specific expressed-sequence-tag (EST) libraries and detailed database annotations, we find a functional genomic signature of rapid evolution and selective constraint shared between mammalian and nematode lineages despite their extensive morphological and ecological differences and distant common ancestry. In both phyla, we find evidence of accelerated evolution among components of molecular systems involved in coevolutionary change. In mammals, lineage-specific fast evolving genes include those involved in reproduction, immunity, and possibly, maternal-fetal conflict. Likelihood ratio tests provide evidence for positive selection in these rapidly evolving functional categories in mammals. In contrast, slowly evolving genes, in terms of amino acid or insertion/deletion (indel) change, in both phyla are involved in core molecular processes such as transcription, translation, and protein transport. Thus, strong purifying selection appears to act on the same core cellular processes in both mammalian and nematode lineages, whereas positive and/or relaxed selection acts on different biological processes in each lineage.

Amino Acid Substitution↗

Physiological genomics of Escherichia coli protein families.

The well-researched Escherichia coli genome offers the opportunity to explore the value of using protein families within a single organism to enrich functional annotation procedures and to study mechanisms of protein evolution. Having identified multimodular proteins resulting from gene fusion, and treated each module as a separate protein, nonoverlapping sequence-similar families in E. coli could be assembled. Of 3,902 proteins of length 100 residues or more, 2,415 clustered into 609 protein families. The relatedness of function among members of each family was dissected in detail. Data on paralogous protein families provides valuable information in attributing putative function to unknown genes, supplementing existing function annotation. Enzymes, transporters, and regulators represent the three major types of proteins in E. coli. They are shown to have distinctive patterns in gene duplication and divergence and gene fusion, suggesting that details of protein evolution have been different for genes in these categories. Data for the complete list of paralogous protein families and updated functional annotation for E. coli K-12 are accessible in GenProtEC (http://genprotec.mbl.edu).

Carrier Proteins↗

Transcriptomic responses of Porphyrophora sophorae larvae during licorice root colonization reveal coordinated remodeling of translation, mitochondrial energy metabolism and defense-related genes.

BACKGROUND: Porphyrophora sophorae is a subterranean piercing-sucking scale insect that damages licorice (Glycyrrhiza uralensis) roots, but the molecular responses associated with larval root colonization remain insufficiently defined. METHODS: We compared non-parasitic larvae (NP) and root-colonizing larvae (RC) using six RNA-seq libraries, de novo transcriptome assembly, DESeq2-based differential expression analysis, GO/KEGG enrichment, annotation-based candidate gene screening, and RT-qPCR validation of selected genes. RESULTS: Sequencing yielded 260.91 million clean reads, and de novo assembly produced 60,794 non-redundant transcripts. DESeq2 identified 703 FDR-significant DEGs, including 49 upregulated and 654 downregulated genes in RC larvae. Upregulated genes were mainly associated with translation- and ribosome-related processes, whereas downregulated genes were enriched in mitochondrial, oxidation-reduction, energy metabolism, and oxidative phosphorylation-related functions. Annotation-based screening identified 75 FDR-significant candidate genes associated with chemosensation, defense-related responses, and energy metabolism, with mitochondrial energy metabolism-related genes forming the largest module. RT-qPCR validation based on the raw Ct data showed concordant expression directions for ten selected transcript targets. CONCLUSIONS: Root colonization in P. sophorae larvae was associated with coordinated transcriptional remodeling involving selective activation of translation-related processes, adjustment of mitochondrial energy metabolism, and changes in defense-related gene expression. These results provide candidate molecular targets for future functional studies of host contact, feeding establishment, and physiological adjustment in this subterranean scale insect.

Animals↗

BASys: a web server for automated bacterial genome annotation.

BASys (Bacterial Annotation System) is a web server that supports automated, in-depth annotation of bacterial genomic (chromosomal and plasmid) sequences. It accepts raw DNA sequence data and an optional list of gene identification information and provides extensive textual annotation and hyperlinked image output. BASys uses >30 programs to determine approximately 60 annotation subfields for each gene, including gene/protein name, GO function, COG function, possible paralogues and orthologues, molecular weight, isoelectric point, operon structure, subcellular localization, signal peptides, transmembrane regions, secondary structure, 3D structure, reactions and pathways. The depth and detail of a BASys annotation matches or exceeds that found in a standard SwissProt entry. BASys also generates colorful, clickable and fully zoomable maps of each query chromosome to permit rapid navigation and detailed visual analysis of all resulting gene annotations. The textual annotations and images that are provided by BASys can be generated in approximately 24 h for an average bacterial chromosome (5 Mb). BASys annotations may be viewed and downloaded anonymously or through a password protected access system. The BASys server and databases can also be downloaded and run locally. BASys is accessible at http://wishart.biology.ualberta.ca/basys.

Chromosomes, Bacterial↗

Contribution of three bile-associated loci, bsh, pva, and btlB, to gastrointestinal persistence and bile tolerance of Listeria monocytogenes.

Listeria monocytogenes must resist the deleterious actions of bile in order to infect and subsequently colonize the human gastrointestinal tract. The molecular mechanisms used by the bacterium to resist bile and the influence of bile on pathogenesis are as yet largely unexplored. This study describes the analysis of three genes--bsh, pva, and btlB--previously annotated as bile-associated loci in the sequenced L. monocytogenes EGDe genome (lmo2067, lmo0446, and lmo0754, respectively). Analysis of deletion mutants revealed a role for all three genes in resisting the acute toxicity of bile and bile salts, particularly glycoconjugated bile salts at low pH. Mutants were unaffected in the other stress responses examined (acid, salt, and detergents). Bile hydrolysis assays demonstrate that L. monocytogenes possesses only one bile salt hydrolase gene, namely, bsh. Transcriptional analyses and activity assays revealed that, although it is regulated by both PrfA and sigma(B), the latter appears to play the greater role in modulating bsh expression. In addition to being incapable of bile hydrolysis, a sigB mutant was shown to be exquisitely sensitive to bile salts. Furthermore, increased expression of sigB was detected under anaerobic conditions and during murine infection. A gene previously annotated as a possible penicillin V amidase (pva) or bile salt hydrolase was shown to be required for resistance to penicillin V but not penicillin G but did not demonstrate a role in bile hydrolysis. Finally, animal (murine) studies revealed an important role for both bsh and btlB in the intestinal persistence of L. monocytogenes.

Animals↗

Identification of candidate variants in plasma associated with early versus late disease progression under anti-PD-1 therapy in metastatic NSCLC.

BACKGROUND: Immune checkpoint inhibitors (ICIs), including anti-programmed cell death protein 1 (anti-PD-1) antibodies, have significantly improved outcomes in patients with metastatic non-small cell lung cancer (mNSCLC). However, substantial heterogeneity exists in clinical benefit, with some patients exhibiting early progression (EP) and others late progression (LP). To date, no biomarkers of EP versus LP disease have been implemented in clinical practice. Circulating tumor DNA (ctDNA) analysis represents a minimally invasive strategy for identifying such biomarkers. In this proof-of-concept study, we evaluated the performance of the TruSight Oncology 500 ctDNA (TSO500 ctDNA) panel and explored its feasibility to identify candidate variants associated with early and late disease progression under anti-PD-1 therapy. METHODS: Baseline ctDNA from eight mNSCLC patients treated with pembrolizumab was extracted and sequenced using the TSO500 ctDNA assay, a 523-gene targeted next-generation sequencing panel. Patients were classified according to their response as LP or EP. Variant calling was performed using the DRAGEN Bio-IT platform, and variants were annotated and clinically interpreted using the Clinical Genomics Workspace (CGW; PierianDx) according to Association for Molecular Pathology (AMP)/American Society of Clinical Oncology (ASCO)/College of American Pathologists (CAP) guidelines. Survival outcomes were assessed using Kaplan-Meier and log-rank tests. Performance of ctDNA variants was evaluated using receiver operating characteristic (ROC) curve analysis, and multi-gene models were assessed using leave-one-out cross-validation with penalized logistic regression. RESULTS: All patients harbored detectable variants, including SNVs (100%), MNVs (87.5%), deletions (75%), and insertions (62.5%). Tier I variants were identified in 37.5% of patients, while all cases showed tier II and multiple tier III alterations. TP53 variants were associated with poorer outcomes under anti-PD-1 therapy. Individual gene alterations in TP53, ERBB3, SMC1A or LATS1 showed moderate discriminatory performance between LP and EP patients; however, combination of mutated genes improved apparent discrimination. Notably, specific two-gene combinations (SMC1A + LATS1 or ERBB3 + LATS1) showed the highest discriminatory performance between LP and EP patients in this exploratory cohort. CONCLUSIONS: This study demonstrates the feasibility and analytical performance of the TSO500 ctDNA panel and provides hypothesis-generating evidence that plasma gene variants may be useful to evaluate early versus late disease progression in patients with mNSCLC receiving immunotherapy.

TruSight Oncology 500↗

Tempo and mode of ERV-K evolution in human and chimpanzee genomes.

Several families of endogenous retrovirus (ERV) exist in copious numbers in the genomes of primate species. Therefore, we undertook a systematic search for endogenous retrovirus sequences from the ERV-K family, comparing across both human (Homo sapiens) and chimpanzee (Pan troglodytes) genomes. Using conserved motifs of the ERV-K as query we identified and characterized 76 complete ERV-K elements, 54 in human (HERV-K), 34 of which were described previously, and 21 in the chimpanzee (CERV-K). Phylogenetic analysis using coding regions and LTRs showed the existence of two main branches. Group I was the most heterogeneous and had an average integration time of 18.3 MYBP (million years before present), using rates ranging from 1.5 to 4.0 x 10(-9) s/s/y (substitution per site per year). Group O/N integrated around 19.4 MYBP and nested Group N integrated about 14 MYBP. We found evidence for strong positive selection on the gag, pol and env coding regions and for A/T hypermutation. Our data suggest that the endogenous elements were possibly involved in chromosomal rearrangements and retained a great deal of information from their active stage, most likely as a consequence of host interactions. This study also contributes to the annotation effort of both human and chimpanzee genomes.

Animals↗

The Ligand Gated Ion Channel database: an example of a sequence database in neuroscience.

Multiple comparisons of receptor sequences, or receptor subunit sequences, has proved to be an invaluable tool in modern pharmacological investigations. Although of outstanding importance, general sequence databases suffer from several imperfections due to their size and their non-specificity. Room therefore exists for expert-maintained databases of restricted focus, where knowledge of the research field helps to filter the huge amount of data generated. Accordingly, neuroscientists have designed databases covering several types of proteins, in particular receptors for neurotransmitters. Ligand-gated ion channels are oligomeric transmembrane proteins involved in the fast response to neurotransmitters. All these receptors are formed by the assembly of homologous subunits, and an unexpected wealth of genes coding for these subunits has been revealed during the last two decades. The Ligand Gated Ion Channel database (LGICdb) has been developed to handle this growing body of information. The database aims to provide only one entry for each gene, containing annotated nucleic acid and protein sequences.

Amino Acid Sequence↗

An efficient algorithm for large-scale detection of protein families.

Detection of protein families in large databases is one of the principal research objectives in structural and functional genomics. Protein family classification can significantly contribute to the delineation of functional diversity of homologous proteins, the prediction of function based on domain architecture or the presence of sequence motifs as well as comparative genomics, providing valuable evolutionary insights. We present a novel approach called TRIBE-MCL for rapid and accurate clustering of protein sequences into families. The method relies on the Markov cluster (MCL) algorithm for the assignment of proteins into families based on precomputed sequence similarity information. This novel approach does not suffer from the problems that normally hinder other protein sequence clustering algorithms, such as the presence of multi-domain proteins, promiscuous domains and fragmented proteins. The method has been rigorously tested and validated on a number of very large databases, including SwissProt, InterPro, SCOP and the draft human genome. Our results indicate that the method is ideally suited to the rapid and accurate detection of protein families on a large scale. The method has been used to detect and categorise protein families within the draft human genome and the resulting families have been used to annotate a large proportion of human proteins.

Algorithms↗

Compilation of mRNA polyadenylation signals in Arabidopsis revealed a new signal element and potential secondary structures.

Using a novel program, SignalSleuth, and a database containing authenticated polyadenylation [poly(A)] sites, we analyzed the composition of mRNA poly(A) signals in Arabidopsis (Arabidopsis thaliana), and reevaluated previously described cis-elements within the 3'-untranslated (UTR) regions, including near upstream elements and far upstream elements. As predicted, there are absences of high-consensus signal patterns. The AAUAAA signal topped the near upstream elements patterns and was found within the predicted location to only approximately 10% of 3'-UTRs. More importantly, we identified a new set, named cleavage elements, of poly(A) signals flanking both sides of the cleavage site. These cis-elements were not previously revealed by conventional mutagenesis and are contemplated as a cluster of signals for cleavage site recognition. Moreover, a single-nucleotide profile scan on the 3'-UTR regions unveiled a distinct arrangement of alternate stretches of U and A nucleotides, which led to a prediction of the formation of secondary structures. Using an RNA secondary structure prediction program, mFold, we identified three main types of secondary structures on the sequences analyzed. Surprisingly, these observed secondary structures were all interrupted in previously constructed mutations in these regions. These results will enable us to revise the current model of plant poly(A) signals and to develop tools to predict 3'-ends for gene annotation.

3' Untranslated Regions↗

Accurate and scalable identification of functional sites by evolutionary tracing.

A common difficulty in post genomics biology is that large-scale techniques of data collection often strip away information on the biological context of these data. The result is a massive number of disconnected observations on sequence, structure, and function from which underlying patterns and biological meaning are obscured. One solution is to build computational filters that pick out sufficiently few facts, relevant to a query, that their relationship is immediately apparent and experimentally testable. Typically, these filters rely on mathematics and statistics, and on first principles from physics and chemistry. We show here that evolution itself can be used to filter sequence and structure data in order to identify evolutionarily important amino acids. A general property of these residues is that they form clusters in native protein structures and point to regions where mutations have the greatest biological impact. The result is an accurate method of functional site annotation that is scalable for structural proteomics.

Amino Acid Sequence↗

Malaria and the red blood cell membrane.

Malaria is the most serious and widespread parasitic disease of humans and is arguably the commonest disease of red blood cells (RBCs). Malaria has exerted a powerful effect on human evolution and selection for resistance has led to the appearance and persistence of a number of inherited diseases. After parasite invasion, RBCs are progressively and dramatically modified. New structures appear inside the RBC and novel parasite proteins are exported to the erythrocyte cytoplasm and membrane skeleton. Radical biochemical, morphological, and rheological alterations manifest as increased membrane rigidity, reduced cell deformability, and greater adhesiveness for the vascular endothelium and other blood cells. Numerous protein-protein interactions between the malaria-parasite and the host RBC are important for many aspects of parasite biology and the pathogenesis of malaria. In addition, there are many other parasite proteins located within the infected red cell and at the membrane skeleton, for which no precise functional roles have yet been elucidated. Sequencing and annotation of the complete genome of Plasmodium falciparum, the production of proteomic and transcriptomic profiles of parasites, and the development of a transfection system for the asexual stage of the parasite are all recent achievements that should advance understanding of the molecular mechanisms that underlie the parasite-induced functional alterations in red cells.

Animals↗

The human biological resource centres network. A key infrastructure for biomedical research in Europe.

Continuing advances in life sciences and medical research are the result of the remarkable achievements of cell and molecular biology. In the post-sequencing era, the quality of the huge amount of data continuously generated by biotechnology, i.e. genomics, proteomics and high-throughput screening, depends on the quality assurance and the traceability of the original biological materials, including the annotations linked to these materials. Thus, biological resource centres are key infrastructures supporting biotechnology, bioprocessing and the development of new approaches in the prevention, diagnosis and treatment of diseases.

Biomedical Research↗

Experimental validation of novel genes predicted in the un-annotated regions of the Arabidopsis genome.

BACKGROUND: Several lines of evidence support the existence of novel genes and other transcribed units which have not yet been annotated in the Arabidopsis genome. Two gene prediction programs which make use of comparative genomic analysis, Twinscan and EuGene, have recently been deployed on the Arabidopsis genome. The ability of these programs to make use of sequence data from other species has allowed both Twinscan and EuGene to predict over 1000 genes that are intergenic with respect to the most recent annotation release. A high throughput RACE pipeline was utilized in an attempt to verify the structure and expression of these novel genes. RESULTS: 1,071 un-annotated loci were targeted by RACE, and full length sequence coverage was obtained for 35% of the targeted genes. We have verified the structure and expression of 378 genes that were not present within the most recent release of the Arabidopsis genome annotation. These 378 genes represent a structurally diverse set of transcripts and encode a functionally diverse set of proteins. CONCLUSION: We have investigated the accuracy of the Twinscan and EuGene gene prediction programs and found them to be reliable predictors of gene structure in Arabidopsis. Several hundred previously un-annotated genes were validated by this work. Based upon this information derived from these efforts it is likely that the Arabidopsis genome annotation continues to overlook several hundred protein coding genes.

Arabidopsis↗

Microarray expression analysis of meiosis and microsporogenesis in hexaploid bread wheat.

BACKGROUND: Our understanding of the mechanisms that govern the cellular process of meiosis is limited in higher plants with polyploid genomes. Bread wheat is an allohexaploid that behaves as a diploid during meiosis. Chromosome pairing is restricted to homologous chromosomes despite the presence of homoeologues in the nucleus. The importance of wheat as a crop and the extensive use of wild wheat relatives in breeding programs has prompted many years of cytogenetic and genetic research to develop an understanding of the control of chromosome pairing and recombination. The rapid advance of biochemical and molecular information on meiosis in model organisms such as yeast provides new opportunities to investigate the molecular basis of chromosome pairing control in wheat. However, building the link between the model and wheat requires points of data contact. RESULTS: We report here a large-scale transcriptomics study using the Affymetrix wheat GeneChip(R) aimed at providing this link between wheat and model systems and at identifying early meiotic genes. Analysis of the microarray data identified 1,350 transcripts temporally-regulated during the early stages of meiosis. Expression profiles with annotated transcript functions including chromatin condensation, synaptonemal complex formation, recombination and fertility were identified. From the 1,350 transcripts, 30 displayed at least an eight-fold expression change between and including pre-meiosis and telophase II, with more than 50% of these having no similarities to known sequences in NCBI and TIGR databases. CONCLUSION: This resource is now available to support research into the molecular basis of pairing and recombination control in the complex polyploid, wheat.

Chromosome Pairing↗