Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

Chromosome-level genome assembly of bivalve mollusk, Xishishe Coelomactra antiquata.

Coelomactra antiquata, a significant marine economic shellfish in China, is experiencing a natural population decline due to habitat destruction and overfishing, making the restoration and conservation of its natural resources an urgent priority. This study provides a high - quality chromosome - level genome assembly for C. antiquata, created by PacBio and Hi - C sequencing and resulting in a 19 - chromosome map. The assembly encompasses a genome size of 807.31 Mb, with a contig N50 of 17.35 Mb and a scaffold N50 of 42.90 Mb. A total of 28,070 protein - coding genes were identified, 25,959 of which were functionally annotated. Overall, this study offers a chromosome - level genome for C. antiquata that is highly continuous and complete, providing an indispensable resource for subsequent molecular and genetic studies of this species.

Animals↗

Chromosomal-level genome assembly of Trypanosoma carassii, the etiologic agent of a recent outbreak of trypanosomiasis in cage-cultured large yellow croaker (Larimichthys crocea) in China.

Trypanosoma carassii, a typical freshwater fish trypanosome, has recently been identified as the etiological agent of a trypanosomiasis outbreak in cage-cultured large yellow croaker (Larimichthys crocea) in China and has been designated as T. c. larimichthys. To date, publicly available genomic data for trypanosomes have been limited to terrestrial species, particularly those of medical importance. Here, we present a chromosome-level genome assembly of T. carassii, the first genome of an aquatic trypanosome, generated using PacBio HiFi long-read sequencing and Hi-C scaffolding technologies. A preliminary genome survey based on Illumina sequencing data estimated the genome size at 56.38 Mb with a heterozygosity of 1.17%. The final assembled genome spans 48.55 Mb, with contig N50 and scaffold N50 values of 139.15 Kb, and achieves 100.00% BUSCO completeness. Hi-C data resolved the assembly into 34 chromosomes and 9 unanchored scaffolds. Repetitive elements account for 53.29% of the genome (approximately 25.87 Mb). A total of 11,584 protein-coding genes were predicted, 95.36% of which were functionally annotated. Synonymous substitution rates analysis of paralogous genes indicates a recent burst of gene duplication, which likely corresponds to a whole-genome duplications. This high-quality genome assembly provides invaluable resources for understanding the evolution and host adaptation of aquatic trypanosomes.

Animals↗

Chromosome-level genome assembly of a cosmopolitan marine harmful algal bloom diatom species Chaetoceros socialis (Chaetocerotaceae).

Chaetoceros socialis is a cosmopolitan diatom species that is crucial for maintaining marine ecosystem structure and driving elemental cycles. C. socialis can form harmful algal blooms (HABs) that may cause a negative impact on the marine ecosystems. Whole-genome information for C. socialis is still unavailable, which may hinder more targeted studies on its ecological adaptive responses and evolutionary drivers. To address this gap, we employed cutting-edge genomic technologies including PacBio single-molecule real-time (SMRT) sequencing and high-throughput chromatin conformation capture (Hi-C) to achieve the first chromosome-level genome assembly of C. socialis. The assembled genome is 60.22 Mb in size with a scaffold N50 of 7.81 Mb and has been anchored to eight pseudochromosomes. A total of 13,378 protein-coding genes were predicted, of which 12,069 (90.22%) were functionally annotated. This high-quality genomic resource provides a fundamental data platform for systematically elucidating the ecological adaptation mechanisms of C. socialis.

Chromosomes↗

A high-quality chromosome-level genome assembly of apple of Peru (Nicandra physalodes).

Nicandra physalodes, a member of the Solanaceae family, is known for its medicinal potential and strong natural insect-repellent properties, which are mainly attributed to its bioactive withanolides and alkaloids. Despite its ecological and pharmacological significance, genomic information for this species has remained limited. Here, we generated a chromosome-level reference genome for N. physalodes based on PacBio high-fidelity (HiFi) long-read sequencing and Hi-C scaffolding. The assembled genome is 933.97 Mb in size, with a contig N50 of 87.37 Mb, and 99.95% (933.54 Mb) of the sequences anchored to 10 pseudochromosomes. Repetitive elements account for 73.06% of the genome, and 27,925 protein-coding genes were predicted, 97.81% of which were functionally annotated. This genomic resource provides a valuable foundation for investigating the genetic basis of specialized metabolite biosynthesis, insect resistance, and environmental adaptation in N. physalodes, as well as for comparative studies within the Solanaceae family.

Genome, Plant↗

A chromosome-level assembly of the alpine snow alga Chloromonas typhlos.

Chloromonas typhlos is a cosmopolitan alpine snow alga distributed across continents, and its blooming accelerates snow melting by decreasing the amount of snow albedo. To elucidate the genetic traits underlying the adaptation of C. typhlos to the alpine habitat, we combined PacBio sequencing and Hi-C to generate a high-quality chromosome-level genome assembly (contig N50: 1.29 Mb; scaffold N50: 7.23 Mb) with 31 chromosomes and a genome size of 200.86 Mb. Repetitive elements constituted 11.05% of the genome, and 16,133 protein-coding genes were predicted, of which 82% were functionally annotated. This study provides a set of omics resources both for snow algae and the genus Chloromonas.

Snow↗

Genomic and phenotypic characterization of Klebsiella pneumoniae phage KP Ø1: a novel lytic Slopekvirus targeting uropathogenic multidrug-resistant Klebsiella pneumoniae.

The rise of multidrug-resistant (MDR) uropathogenic gram-negative bacteria (GNB) necessitates the development of alternative therapeutic strategies. This study aimed to isolate, phenotypically characterize, and perform whole-genome sequencing of the bacteriophage demonstrating the broadest host range against MDR uropathogens. Fifty MDR GNB isolates were screened for lytic phages. The most promising candidate, Klebsiella pneumoniae phage KP Ø1, was characterized using plaque assay, Transmission Electron Microscopy (TEM), and pH/thermal stability testing. Genomic characterization was performed via whole-genome sequencing (WGS), with functional annotation and lifestyle prediction using PhaBOX and PhageScope software. Klebsiella pneumoniae was the most prevalent MDR uropathogen. Klebsiella pneumoniae phage KP Ø1 exhibited a 50% host range and high lytic titer (10⁸ PFU/mL). TEM revealed an icosahedral head and short contractile tail. Genomic characterization by WGS revealed that Klebsiella pneumoniae phage KP Ø1 possesses a 174,591 bp double-stranded deoxyribonucleic acid (dsDNA) genome containing 274 predicted open reading frames (ORFs). No lysogeny-related genes, toxins, or antibiotic resistance markers were detected, confirming its strictly lytic nature and supporting its potential as a candidate for phage therapy applications. The phage remained stable (10⁸ PFU/mL) across temperatures of - 20 °C to 50 °C; supporting its suitability for long-term biobanking and suggesting potential activity at physiological temperature, and across a pH range of 7-9. Klebsiella pneumoniae phage KP Ø1 is a novel, obligately lytic Slopekvirus whose genomic architecture, stability profile, and absence of lysogeny-associated, virulence, and antimicrobial resistance genes ( AMR) collectively support its candidacy for further preclinical evaluation as a phage therapy agent against uropathogenic MDR Klebsiella pneumoniae.

Klebsiella pneumoniae↗

The apoptosis database.

The apoptosis database is a public resource for researchers and students interested in the molecular biology of apoptosis. The resource provides functional annotation, literature references, diagrams/images, and alternative nomenclatures on a set of proteins having 'apoptotic domains'. These are the distinctive domains that are often, if not exclusively, found in proteins involved in apoptosis. The initial choice of proteins to be included is defined by apoptosis experts and bioinformatics tools. Users can browse through the web accessible lists of domains, proteins containing these domains and their associated homologs. The database can also be searched by sequence homology using basic local alignment search tool, text word matches of the annotation, and identifiers for specific records. The resource is available at http://www.apoptosis-db.org and is updated on a regular basis.

Animals↗

Robust functional gene validation by adenoviral vectors: one-step Escherichia coli-derived recombinant adenoviral genome construction.

We describe here a clonal approach for efficient and robust construction of recombinant adenoviral genomes that holds certain advantages over existing approaches. Transgenes of interest are cloned into a small, conditionally replicating plasmid containing the left end of a recombinant adenoviral genome, encompassing pIX coding regions. Transformation of this plasmid into recombination-competent Escherichia coli bearing a plasmid containing the right end of a recombinant adenoviral genome, commencing from pIX coding regions, yields a stable co-integrated plasmid encoding a full adenoviral genome, by virtue of shared homology in pIX coding regions contained in both plasmids. The recombination process yielding the full adenoviral plasmid requires only one step, and always results in the formation of only the desired recombinant adenoviral genome. Thus, no screening is required to identify the correct plasmid encoding the desired recombinant adenoviral genome. In addition, the plasmid encoding the right-hand side of the adenoviral genome is itself incapable of producing contaminating adenovirus. We have successfully employed this approach to generate over 200 recombinant adenoviruses, obtaining only the desired recombinant adenoviral species each time. The process is amenable to medium-to-high-throughput parallel construction of adenoviral genomes, and as such should aid efforts aimed towards high-throughput functional annotation of therapeutic gene targets, which aim to leverage the benefits of adenoviruses as gene delivery and expression vectors.

Adenoviridae↗

Gene expression profiles of human proximal tubular epithelial cells in proteinuric nephropathies.

In kidney disease renal proximal tubular epithelial cells (RPTEC) actively contribute to the progression of tubulointerstitial fibrosis by mediating both an inflammatory response and via epithelial-to-mesenchymal transition. Using laser capture microdissection we specifically isolated RPTEC from cryosections of the healthy parts of kidneys removed owing to renal cell carcinoma and from kidney biopsies from patients with proteinuric nephropathies. RNA was extracted and hybridized to complementary DNA microarrays after linear RNA amplification. Statistical analysis identified 168 unique genes with known gene ontology association, which separated patients from controls. Besides distinct alterations in signal-transduction pathways (e.g. Wnt signalling), functional annotation revealed a significant upregulation of genes involved in cell proliferation and cell cycle control (like insulin-like growth factor 1 or cell division cycle 34), cell differentiation (e.g. bone morphogenetic protein 7), immune response, intracellular transport and metabolism in RPTEC from patients. On the contrary we found differential expression of a number of genes responsible for cell adhesion (like BH-protocadherin) with a marked downregulation of most of these transcripts. In summary, our results obtained from RPTEC revealed a differential regulation of genes, which are likely to be involved in either pro-fibrotic or tubulo-protective mechanisms in proteinuric patients at an early stage of kidney disease.

Aged↗

Meta-analysis of microarray data on pancreatic cancer defines a set of commonly dysregulated genes.

Pancreatic ductal adenocarcinoma is the eighth most common cancer with the lowest overall 5-year relative survival rate of any tumor type today. Expression profiling using microarrays has been widely used to identify genes associated with pancreatic cancer development. To extract maximum value from the available gene expression data, we applied a meta-analysis to search for commonly differentially expressed genes in pancreatic ductal adenocarcinoma. We obtained data sets from four different gene expression studies on pancreatic cancer. We selected a consensus set of 2984 genes measured in all four studies and applied a meta-analysis approach to evaluate the combined data. Of the genes identified as differentially expressed, several were validated using RT-PCR and immunohistochemistry. Additionally, we used a class discovery algorithm to identify a gene expression signature. Our meta-analysis revealed that the pancreatic cancer gene expression data sets shared a significant number of up- and downregulated genes, independent of the technology used. This interstudy crossvalidation approach generated a set of 568 genes that were consistently and significantly dysregulated in pancreatic cancer. Of these, 364 (64.1%) were upregulated and 204 (35.9%) were downregulated in pancreatic cancer. Only 127 (22%) were described in the published individual analyses. Functional annotation of the genes revealed that genes presumably associated with the cell adhesion-mediated drug resistance pathway are frequently overexpressed in pancreatic cancer. Meta-analysis is an important tool for the identification and validation of differentially expressed genes. These could represent good candidates for novel diagnostic and therapeutic approaches to pancreatic cancer.

Adenocarcinoma↗

Identification of novel tumour-associated genes differentially expressed in the process of squamous cell cancer development.

Chemically induced mouse skin carcinogenesis represents the most extensively utilized animal model to unravel the multistage nature of tumour development and to design novel therapeutic concepts of human epithelial neoplasia. We combined this tumour model with comprehensive gene expression analysis and could identify a large set of novel tumour-associated genes that have not been associated with epithelial skin cancer development yet. Expression data of selected genes were confirmed by semiquantitative and quantitative RT-PCR as well as in situ hybridization and immunofluorescence analysis on mouse tumour sections. Enhanced expression of genes identified in our screen was also demonstrated in mouse keratinocyte cell lines that form tumours in vivo. Self-organizing map clustering was performed to identify different kinetics of gene expression and coregulation during skin cancer progression. Detailed analysis of differential expressed genes according to their functional annotation confirmed the involvement of several biological processes, such as regulation of cell cycle, apoptosis, extracellular proteolysis and cell adhesion, during skin malignancy. Finally, we detected high transcript levels of ANXA1, LCN2 and S100A8 as well as reduced levels for NDR2 protein in human skin tumour specimens demonstrating that tumour-associated genes identified in the chemically induced tumour model might be of great relevance for the understanding of human epithelial malignancies as well.

Animals↗

Head and neck squamous cell carcinoma transcriptome analysis by comprehensive validated differential display.

Head and neck squamous cell carcinoma (HNSCC) is common worldwide and is associated with a poor rate of survival. Identification of new markers and therapeutic targets, and understanding the complex transformation process, will require a comprehensive description of genome expression, that can only be achieved by combining different methodologies. We report here the HNSCC transcriptome that was determined by exhaustive differential display (DD) analysis coupled with validation by different methods on the same patient samples. The resulting 820 nonredundant sequences were analysed by high throughput bioinformatics analysis. Human proteins were identified for 73% (596) of the DD sequences. A large proportion (>50%) of the remaining unassigned sequences match ESTs (expressed sequence tags) from human tumours. For the functionally annotated proteins, there is significant enrichment for relevant biological processes, including cell motility, protein biosynthesis, stress and immune responses, cell death, cell cycle, cell proliferation and/or maintenance and transport. Three of the novel proteins (TMEM16A, PHLDB2 and ARHGAP21) were analysed further to show that they have the potential to be developed as therapeutic targets.

Amino Acid Sequence↗

A new paradigm for drug discovery: integrating clinical, genetic, genomic and molecular phenotype data to identify drug targets.

Application of statistical genetics approaches to variations in mRNA transcript abundances in segregating populations can be used to identify genes and pathways associated with common human diseases. The combination of this genetic information with gene expression and clinical trait data can also be used to identify subtypes of a disease and the genetic loci specific to each subtype. Here we highlight results from some of our recent work in this area and further explore the many possibilities that exist in employing a more comprehensive genetics and functional genomics approach to the functional annotation of genomes, and in applying such methods to the validation of targets for complex traits in the drug discovery process.

Animals↗

Primary and secondary metabolism, and post-translational protein modifications, as portrayed by proteomic analysis of Streptomyces coelicolor.

The newly sequenced genome of Streptomyces coelicolor is estimated to encode 7825 theoretical proteins. We have mapped approximately 10% of the theoretical proteome experimentally using two-dimensional gel electrophoresis and matrix-assisted laser desorption ionization time-of-flight (MALDI-TOF) mass spectrometry. Products from 770 different genes were identified, and the types of proteins represented are discussed in terms of their annotated functional classes. An average of 1.2 proteins per gene was observed, indicating extensive post-translational regulation. Examples of modification by N-acetylation, adenylylation and proteolytic processing were characterized using mass spectrometry. Proteins from both primary and certain secondary metabolic pathways are strongly represented on the map, and a number of these enzymes were identified at more than one two-dimensional gel location. Post-translational modification mechanisms may therefore play a significant role in the regulation of these pathways. Unexpectedly, one of the enzymes for synthesis of the actinorhodin polyketide antibiotic appears to be located outside the cytoplasmic compartment, within the cell wall matrix. Of 20 gene clusters encoding enzymes characteristic of secondary metabolism, eight are represented on the proteome map, including three that specify the production of novel metabolites. This information will be valuable in the characterization of the new metabolites.

Acetylation↗

Microbial disease in humans: A genomic perspective.

The approach of whole-genome shotgun sequencing coupled with the availability of computational algorithms to facilitate the assembly, gene prediction, and functional annotation of entire genomes has sparked a revolution in our understanding of the biology of free-living organisms. More than 40 bacterial genomes have been sequenced to date, of which several are important human pathogens. The capacity to sequence and assemble entire genomes of bacteria, pathogenic protozoans, and fungi in a rapid and cost-effective way has energized every aspect of microbial science. Comparative genome analysis allows us to dissect the evolutionary forces at work and provides insights into adaptations of microbes to their unique ecological niches. Factors that shape host-pathogen interactions and their outcomes include genetic polymorphisms in the microbial pathogen and host, both of which can impact on microbial virulence or host immune responses to infection. The availability of the genome sequence of entire organisms, together with the use of high-throughput sequence-based genomic technologies to define microbial and host physiological states, provides the unparalleled opportunity to better define clinical outcomes in the field of infectious diseases. There is one overarching lesson: completion of the genomic sequence of any species answers many questions, while at the same time it invites totally new questions.

Animals↗

Protein structure prediction for the male-specific region of the human Y chromosome.

The complete sequence of the male-specific region of the human Y chromosome (MSY) has been determined recently; however, detailed characterization for many of its encoded proteins still remains to be done. We applied state-of-the-art protein structure prediction methods to all 27 distinct MSY-encoded proteins to provide better understanding of their biological functions and their mechanisms of action at the molecular level. The results of such large-scale structure-functional annotation provide a comprehensive view of the MSY proteome, shedding light on MSY-related processes. We found that, in total, at least 60 domains are encoded by 27 distinct MSY genes, of which 42 (70%) were reliably mapped to currently known structures. The most challenging predictions include the unexpected but confident 3D structure assignments for three domains identified here encoded by the USP9Y, UTY, and BPY2 genes. The domains with unknown 3D structures that are not predictable with currently available theoretical methods are established as primary targets for crystallographic or NMR studies. The data presented here set up the basis for additional scientific discoveries in human biology of the Y chromosome, which plays a fundamental role in sex determination.

Amino Acid Sequence↗

A high-throughput, near-saturating screen for type III effector genes from Pseudomonas syringae.

Pseudomonas syringae strains deliver variable numbers of type III effector proteins into plant cells during infection. These proteins are required for virulence, because strains incapable of delivering them are nonpathogenic. We implemented a whole-genome, high-throughput screen for identifying P. syringae type III effector genes. The screen relied on FACS and an arabinose-inducible hrpL sigma factor to automate the identification and cloning of HrpL-regulated genes. We determined whether candidate genes encode type III effector proteins by creating and testing full-length protein fusions to a reporter called Delta79AvrRpt2 that, when fused to known type III effector proteins, is translocated and elicits a hypersensitive response in leaves of Arabidopsis thaliana expressing the RPS2 plant disease resistance protein. Delta79AvrRpt2 is thus a marker for type III secretion system-dependent translocation, the most critical criterion for defining type III effector proteins. We describe our screen and the collection of type III effector proteins from two pathovars of P. syringae. This stringent functional criteria defined 29 type III proteins from P. syringae pv. tomato, and 19 from P. syringae pv. phaseolicola race 6. Our data provide full functional annotation of the hrpL-dependent type III effector suites from two sequenced P. syringae pathovars and show that type III effector protein suites are highly variable in this pathogen, presumably reflecting the evolutionary selection imposed by the various host plants.

Arabidopsis↗

Predicting ligand-binding function in families of bacterial receptors.

The three-dimensional fold of a new protein sequence can often be inferred directly from sequence homology to a protein of known structure. The function of a new protein sequence is more difficult to predict, however, since homologues can have different molecular and cellular functions. To develop and automate computational methods for determining molecular function, we have analyzed ligand-binding specificity in two related families of binding proteins. One of these families includes Escherichia coli lactose repressor and ribose-binding protein, and the other includes E. coli sulfate- and phosphate-binding proteins. These proteins have similar folds but varying specificity, binding many different small molecules, including mono- and disaccharides, purines, oxyanions, ferric iron, and polyamines. Starting from template structural alignments, alignments of over 90 sequences per family were generated by iterative database searches with hidden Markov models. Phylogenetic trees were made of full-length sequences and of subsets of residues lining the binding cleft, to determine whether subbranches of the trees correlate with ligand-binding preference. Automated analyses of residues in the binding pocket were also used to predict ligand-binding function for many uncharacterized database sequences and to identify specific side chain-ligand contacts in proteins without solved structures. Our results demonstrate the utility of anchoring functional annotation within a protein family context.

Amino Acid Sequence↗