Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Comprehensive genomic profiling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

Gene networks as a tool to understand transcriptional regulation.

Gene regulatory networks, or simply gene networks (GNs), have shown to be a promising approach that the bioinformatics community has been developing for studying regulatory mechanisms in biological systems. GNs are built from the genome-wide high-throughput gene expression data that are often available from DNA microarray experiments. Conceptually, GNs are (un)directed graphs, where the nodes correspond to the genes and a link between a pair of genes denotes a regulatory interaction that occurs at transcriptional level. In the present study, we had two objectives: 1) to develop a framework for GN reconstruction based on a Bayesian network model that captures direct interactions between genes through nonparametric regression with B-splines, and 2) to demonstrate the potential of GNs in the analysis of expression data of a real biological system, the yeast pheromone response pathway. Our framework also included a number of search schemes to learn the network. We present an intuitive notion of GN theory as well as the detailed mathematical foundations of the model. A comprehensive analysis of the consistency of the model when tested with biological data was done through the analysis of the GNs inferred for the yeast pheromone pathway. Our results agree fairly well with what was expected based on the literature, and we developed some hypotheses about this system. Using this analysis, we intended to provide a guide on how GNs can be effectively used to study transcriptional regulation. We also discussed the limitations of GNs and the future direction of network analysis for genomic data. The software is available upon request.

Bayes Theorem↗

A machine learning-derived and functionally validated circadian rhythm signature predicts clinical outcomes and in silico drug sensitivity in colorectal cancer.

BACKGROUND: Colorectal cancer (CRC) displays considerable heterogeneity in clinical outcomes, highlighting the need for reliable prognostic biomarkers. While the aberrant expression of circadian rhythm-related genes has been implicated in cancer pathogenesis, its comprehensive role in CRC progression and predicted therapeutic vulnerabilities remains inadequately characterized. METHODS: Bulk and single-cell RNA-sequencing data were integrated from multiple CRC cohorts. A circadian rhythm signature (CRS) was developed through machine learning algorithms and validated for prognostic value. Comprehensive analyses of tumor microenvironment, genomic alterations, and drug sensitivity were performed. Furthermore, the biological function of the core gene, BHLHE40, was validated in CRC cell lines through CCK-8, EdU, and wound healing assays. RESULTS: Single-cell analysis demonstrated an elevated expression signature of circadian rhythm-related genes in dendritic cells. The optimized CRS, comprising 14 circadian rhythm-related genes, successfully categorized patients into high- and low-risk groups. Patients with a high CRS showed markedly poorer overall survival and computationally inferred immunosuppressive features, including reduced CD8+ T cell infiltration and increased M2 macrophage polarization. Genomic analysis revealed enhanced mutation burden in TP53 and alterations in RTK-RAS/WNT pathways. Notably, in vitro assays confirmed that BHLHE40 is significantly overexpressed in CRC cells. Knockdown of BHLHE40 markedly inhibited tumor cell proliferation and migration. Drug sensitivity profiling identified bexarotene and SMER-3 as potential therapeutic options for high-CRS patients. A nomogram integrating CRS with clinical parameters demonstrated superior predictive accuracy for 1-, 3-, and 5-year survival. CONCLUSIONS: The CRS represents a promising prognostic biomarker that reflects tumor immune status and genomic features, providing valuable insights for personalized treatment strategies in CRC.

Circadian rhythm↗

The rapid identification of Acinetobacter species using Fourier transform infrared spectroscopy.

AIMS: Fourier transform infrared (FT-IR) was used to analyse a selection of Acinetobacter isolates in order to determine if this approach could discriminate readily between the known genomic species of this genus and environmental isolates from activated sludge. METHODS AND RESULTS: FT-IR spectroscopy is a rapid whole-organism fingerprinting method, typically taking only 10 s per sample, and generates 'holistic' biochemical profiles (or 'fingerprints') from biological materials. The cluster analysis produced by FT-IR was compared with previous polyphasic taxonomic studies on these isolates and with 16S-23S rDNA intergenic spacer region (ISR) fingerprinting presented in this paper. FT-IR and 16S-23S rDNA ISR analyses together indicate that some of the Acinetobacter genomic species are particularly heterogeneous and poorly defined, making characterization of the unknown environmental isolates with the genomic species difficult. CONCLUSIONS: Whilst the characterization of the isolates from activated sludge revealed by FT-IR and 16S-23S rDNA ISR were not directly comparable, the dendrogram produced from FT-IR data did correlate well with the outcomes of the other polyphasic taxonomic work. SIGNIFICANCE AND IMPACT OF THE STUDY: We believe it would be advantageous to pursue this approach further and establish a comprehensive database of taxonomically well-defined Acinetobacter species to aid the identification of unknown strains. In this instance, FT-IR may provide the rapid identification method eagerly sought for the routine identification of Acinetobacter isolates from a wide range of environmental sources.

Acinetobacter↗

The Yeast Proteome Database (YPD): a model for the organization and presentation of genome-wide functional data.

The Yeast Proteome Database (YPD) is a model for the organization and presentation of comprehensive protein information. Based on the detailed curation of the scientific literature for the yeast Saccharomyces cerevisiae, YPD contains more than 50 000 annotations lines derived from the review of 8500 research publications. The information concerning each of the approximately 6100 yeast proteins is structured around a convenient one-page format, the Yeast Protein Report, with additional information provided as pop-up windows. Protein classification schema have been revised this year, defining each protein's cellular role, function and pathway, and adding a Functional to the Yeast Protein Report. These changes provide the user with a succinct summary of the protein's function and its place in the biology of the cell, and they enhance the power of YPD Search functions. Precalculated sequence alignments have been added, to provide a crossover point for comparative genomics. The first transcript profiling data has been integrated into the YPD Protein Reports, providing the framework for the presentation of genome-wide functional data. The Yeast Proteome Database can be accessed on the Web at http://www.proteome.com/YPDhome.html

Computational Biology↗

Comprehensive transcriptomic analysis of BjGL1-knockout Brassica juncea: novel insights into leaf trichome formation.

Brassica juncea is a common cruciferous crop, which can be used not only for oil extraction but also as condiments and medicinal materials. It is regarded by both traditional medicine and modern nutrition science as a food with combined dietary and health promoting value. Leaf trichomes are hair-like structures differentiated from epidermal cells and constitute an important barrier against biotic and abiotic stresses, playing a crucial role in enhancing plant resistance and thus possessing significant scientific relevance. In this study, the phenotype and gene editing site of BjA06.GL1 and BjB02.GL1 knockout mustard T1 generation plants were identified. Then, RNA sequencing was performed to compare the leaf transcriptome profiles between gene-edited lines and wild-type plants, with the aim of elucidating the molecular regulatory mechanisms by which BjGL1 controls leaf trichome development and associated biological processes in mustard. The sequencing data showed that, on average, 90.64% of the reads uniquely aligned to the Brassica juncea (Xuecai) reference genome. A total of 4,604 differentially expressed genes were identified in this study. Compared with the gene knockout mutant, 1,831 genes were significantly upregulated and 2,773 genes were downregulated in mustard leaves with trichomes. The differentially expressed genes were mainly enriched in pathways related to cytochrome P450 (CYP), transporters, environmental adaptation, and plant-pathogen interactions. These pathways are closely associated with secondary metabolite biosynthesis, transmembrane transport, and responses to abiotic stress and pathogen defense. qRT-PCR validation confirmed consistent expression trends of trichome regulatory genes screened from transcriptome data. This study provides an important theoretical basis for elucidating molecular mechanisms potentially contributing to trichome formation in mustard.

Mustard Plant↗

BoCaTFBS: a boosted cascade learner to refine the binding sites suggested by ChIP-chip experiments.

Comprehensive mapping of transcription factor binding sites is essential in postgenomic biology. For this, we propose a mining approach combining noisy data from ChIP (chromatin immunoprecipitation)-chip experiments with known binding site patterns. Our method (BoCaTFBS) uses boosted cascades of classifiers for optimum efficiency, in which components are alternating decision trees; it exploits interpositional correlations; and it explicitly integrates massive negative information from ChIP-chip experiments. We applied BoCaTFBS within the ENCODE project and showed that it outperforms many traditional binding site identification methods (for instance, profiles).

Algorithms↗

DARFA: a novel technique for studying differential gene expression and bacterial comparative genomics.

We have developed a powerful method, named differential analysis of restriction fragments amplification (DARFA), which enables researchers to perform comprehensive transcriptome analysis as well as bacterial DNA fingerprinting. The key feature of this novel technique lies within the usage of a type IIS enzyme, Hpy188III, which cleaves cDNA or genomic DNA at a TC/NNGA recognition sequence. Cleavage at this particular site results in the production of a pool of restriction fragments which can be divided into 120 subsets based on the 2-nt 5'-overhang sequence. Each subset of restriction fragments is then selectively amplified by PCR after ligation with a pair of hairpin adaptors containing 2-nt overhangs which are complementary to those in the subset of fragments that are to be analyzed. The results obtained from the analysis of strain-specific and tissue-specific differences using DARFA and further confirmation by DNA sequencing and Northern analysis have demonstrated that the DARFA technique provides a novel tool for expression profiling, as well as bacterial DNA fingerprinting.

Bacteria↗

Sexually dimorphic expression and hormonal responsiveness of steroidogenic Cyp genes during gonadal differentiation in mandarin fish.

Steroid hormones play a pivotal role in fish sex differentiation, yet the dynamic expression patterns of key steroidogenic enzymes during this process remain incompletely characterized. Here, we combined genome-wide identification, time series transcriptomes spanning gonadal development (5-360 days post-hatch), and multiple hormone treatment experiments (17α-methyltestosterone, estrone, and etonogestrel) to investigate the Cyp11, Cyp17, Cyp19, and Cyp21 subfamilies in mandarin fish (Siniperca chuatsi). Seven steroidogenic Cyp genes were identified, showing teleost-specific expansion, with one duplicated pair (cyp17a2 and cyp2u1) exhibiting strong purifying selection. Expression profiling revealed pronounced sexually dimorphic and stage-specific patterns: During female differentiation (20-30 days), cyp19a1a and associated genes were highly expressed, coinciding with ovarian differentiation; during male differentiation (30-60 days), cyp17a2 and related genes were upregulated, aligning with testicular development. Exogenous hormone treatments further demonstrated that these genes are dynamically responsive: cyp19a1a and cyp17a2 were highly responsive to androgenic and progestogenic treatments, and their expression changes correlated closely with gonadal sex reversal phenotypes observed histologically. Collectively, this study provides a comprehensive expression atlas of steroidogenic Cyp genes during gonadal differentiation and identifies key hormonally responsive candidates for sex control in aquaculture.

Animals↗

The molecular similarity landscape of preclinical cancer models to patient tumors.

Selecting appropriate preclinical models is fundamental for translational oncology, yet a large-scale, multi-omic quantitative comparison of their similarity to primary human tumors is lacking. To address this, we integrated transcriptomic, proteomic, and genomic profiles from over 10,000 primary tumors from The Cancer Genome Atlas (TCGA) and the Clinical Proteomic Tumor Analysis Consortium (CPTAC), alongside 4,000 preclinical models. Using a robust computational framework, we revealed a clear hierarchy of transcriptomic and proteomic similarity to patient tumors: with patient-dervied xenografts (PDXs) having greater transcriptomic and proteomic similarity to patient tumors (>) compared with patient-derived organoids (PDOs), which are equal in hierarchy to that of PDX-dervied organoids (PDXOs) > cell lines. We also quantified high molecular conservation (Pearson correlation coefficient = 0.96) across paired in vitro to in vivo platform (organoids to PDX) transitions. Furthermore, genomic analysis demonstrated that whole-exome sequencing (WES) outperforms RNA-seq in detecting DNA variants, and it identified a clonal complexity hierarchy (cell lines > PDXOs > PDXs > PDOs) reflecting the effect of passaging history on intratumor heterogeneity. Ultimately, this study delivers a comprehensive quantitative benchmark, establishing a population-level hierarchy of molecular similarity between preclinical models and primary tumors and providing a data-driven reference for model selection. These findings offer a data-driven framework for selecting models that balance biological representativeness with experimental practicality.

Humans↗

Pfam: a comprehensive database of protein domain families based on seed alignments.

Databases of multiple sequence alignments are a valuable aid to protein sequence classification and analysis. One of the main challenges when constructing such a database is to simultaneously satisfy the conflicting demands of completeness on the one hand and quality of alignment and domain definitions on the other. The latter properties are best dealt with by manual approaches, whereas completeness in practice is only amenable to automatic methods. Herein we present a database based on hidden Markov model profiles (HMMs), which combines high quality and completeness. Our database, Pfam, consists of parts A and B. Pfam-A is curated and contains well-characterized protein domain families with high quality alignments, which are maintained by using manually checked seed alignments and HMMs to find and align all members. Pfam-B contains sequence families that were generated automatically by applying the Domainer algorithm to cluster and align the remaining protein sequences after removal of Pfam-A domains. By using Pfam, a large number of previously unannotated proteins from the Caenorhabditis elegans genome project were classified. We have also identified many novel family memberships in known proteins, including new kazal, Fibronectin type III, and response regulator receiver domains. Pfam-A families have permanent accession numbers and form a library of HMMs available for searching and automatic annotation of new protein sequences.

Amino Acid Sequence↗

Integrated analysis of yeast regulatory sequences for biologically linked clusters of genes.

Dramatic progress in deciphering the regulatory controls in Saccharomyces cerevisiae has been enabled by the fusion of high-throughput genomics technologies with advanced sequence analysis algorithms. Sets of genes likely to function together and with similar expression profiles have been identified in diverse studies. By fusing an advanced pattern recognition algorithm for identification of transcription factor binding sites with a new method for the quantitative comparison of binding properties of transcription factors, we provide an integrated means to move from expression data to biological insights. The Yeast Regulatory Sequence Analysis system, YRSA, combines standard functions with a novel pattern characterization procedure in an intuitive interface designed for use by a broad range of scientists. The features of the system include automated retrieval of user-defined promoter sequences, binding site discovery by pattern recognition, graphical displays of the observed pattern and positions of similar sequences in the specified genes, and comparison of the new pattern against a collection of binding patterns for characterized transcription factors. The comprehensive YRSA system was used to study the regulatory mechanisms of yeast regulons. Analysis of the regulatory controls of a battery of genes induced by DNA damaging agents supports a putative mediating role for the cell-cycle checkpoint regulatory element MCB. YRSA is available at http://yrsa.cgb.ki.se. [YRSA: ancient Scandinavian name meaning old she-bear (Latin Ursus arctos = brown bear/grizzly).]

Algorithms↗

A knowledge-based approach in designing combinatorial or medicinal chemistry libraries for drug discovery. 1. A qualitative and quantitative characterization of known drug databases.

The discovery of various protein/receptor targets from genomic research is expanding rapidly. Along with the automation of organic synthesis and biochemical screening, this is bringing a major change in the whole field of drug discovery research. In the traditional drug discovery process, the industry tests compounds in the thousands. With automated synthesis, the number of compounds to be tested could be in the millions. This two-dimensional expansion will lead to a major demand for resources, unless the chemical libraries are made wisely. The objective of this work is to provide both quantitative and qualitative characterization of known drugs which will help to generate "drug-like" libraries. In this work we analyzed the Comprehensive Medicinal Chemistry (CMC) database and seven different subsets belonging to different classes of drug molecules. These include some central nervous system active drugs and cardiovascular, cancer, inflammation, and infection disease states. A quantitative characterization based on computed physicochemical property profiles such as log P, molar refractivity, molecular weight, and number of atoms as well as a qualitative characterization based on the occurrence of functional groups and important substructures are developed here. For the CMC database, the qualifying range (covering more than 80% of the compounds) of the calculated log P is between -0.4 and 5.6, with an average value of 2.52. For molecular weight, the qualifying range is between 160 and 480, with an average value of 357. For molar refractivity, the qualifying range is between 40 and 130, with an average value of 97. For the total number of atoms, the qualifying range is between 20 and 70, with an average value of 48. Benzene is by far the most abundant substructure in this drug database, slightly more abundant than all the heterocyclic rings combined. Nonaromatic heterocyclic rings are twice as abundant as the aromatic heterocycles. Tertiary aliphatic amines, alcoholic OH and carboxamides are the most abundant functional groups in the drug database. The effective range of physicochemical properties presented here can be used in the design of drug-like combinatorial libraries as well as in developing a more efficient corporate medicinal chemistry library.

Algorithms↗

Comprehensive de novo structure prediction in a systems-biology context for the archaea Halobacterium sp. NRC-1.

BACKGROUND: Large fractions of all fully sequenced genomes code for proteins of unknown function. Annotating these proteins of unknown function remains a critical bottleneck for systems biology and is crucial to understanding the biological relevance of genome-wide changes in mRNA and protein expression, protein-protein and protein-DNA interactions. The work reported here demonstrates that de novo structure prediction is now a viable option for providing general function information for many proteins of unknown function. RESULTS: We have used Rosetta de novo structure prediction to predict three-dimensional structures for 1,185 proteins and protein domains (<150 residues in length) found in Halobacterium NRC-1, a widely studied halophilic archaeon. Predicted structures were searched against the Protein Data Bank to identify fold similarities and extrapolate putative functions. They were analyzed in the context of a predicted association network composed of several sources of functional associations such as: predicted protein interactions, predicted operons, phylogenetic profile similarity and domain fusion. To illustrate this approach, we highlight three cases where our combined procedure has provided novel insights into our understanding of chemotaxis, possible prophage remnants in Halobacterium NRC-1 and archaeal transcriptional regulators. CONCLUSIONS: Simultaneous analysis of the association network, coordinated mRNA level changes in microarray experiments and genome-wide structure prediction has allowed us to glean significant biological insights into the roles of several Halobacterium NRC-1 proteins of previously unknown function, and significantly reduce the number of proteins encoded in the genome of this haloarchaeon for which no annotation is available.

Archaeal Proteins↗

A genomic view of eukaryotic DNA replication.

Recent advances in DNA microarray technology have enabled eukaryotic replication to be studied at whole-chromosome and genome-wide levels. These studies have provided new insights into the mechanisms that influence origin selection and the temporally co-ordinated activation of replication initiation from these sites. Here we describe multiple microarray-based approaches that have been used to study DNA replication in both S. cerevisiae and higher eukaryotes. We have also compiled the data from the yeast microarray-based replication studies to generate a comprehensive list of origins that has been verified in three independent studies. The comprehensive nature of the microarray-based studies has revealed clear connections between chromosome organization and the pattern of replication. For example, in yeast, the centromeric proximal sequences are consistently early replicating and telomeric regions are consistently late replicating. The metazoan studies reveal a recurring theme of gene-dense transcriptionally active regions of the genome replicating before gene-sparse regions. In addition to the insights they have provided already, microarray-based replication assays combined with genetic analysis will provide a powerful new approach to define the mechanisms that regulate replication origin function.

Animals↗

Post-transcriptional control of gene expression: a genome-wide perspective.

Gene expression is regulated at multiple levels, and cells need to integrate and coordinate different layers of control to implement the information in the genome. Post-transcriptional levels of regulation such as transcript turnover and translational control are an integral part of gene expression and might rival the sophistication and importance of transcriptional control. Microarray-based methods are increasingly used to study not only transcription but also global patterns of transcript decay and translation rates in addition to comprehensively identify targets of RNA-binding proteins. Such large-scale analyses have recently provided supplementary and unique insights into gene expression programs. Integration of several different datasets will ultimately lead to a system-wide understanding of the varied and complex mechanisms for gene expression control.

Gene Expression Profiling↗

Molecular surveillance of foodborne bacterial pathogens and resistome in food products from Hong Kong.

Foodborne infections pose an increasing public health challenge worldwide. The problem has been aggravated by the dissemination of antimicrobial resistance genes among zoonotic pathogens, which results in a sharp increase in antibiotic resistance rate recorded among the major foodborne pathogens. To obtain an overview of the extent to which food products purchased in the markets in Hong Kong were contaminated by foodborne pathogens, we collected 95 raw meat samples from wet markets and isolated 236 bacterial strains of various species, with Escherichia coli being the most dominant species (131 strains). Contamination of food products by multiple foodborne pathogens was commonly observed. These include both Gram-positive and Gram-negative bacteria that exhibit various levels of resistance, with some possessing multiple clinically important antibiotic resistance genes. Seventeen bacterial strains of various species isolated from three food samples were comprehensively analysed by the Oxford Nanopore R10.4 technology. Novel conjugative plasmids carrying antimicrobial resistance gene-bearing mobile genetic elements were commonly detectable in the test strains. Some of the plasmids were shown to have originated from other environmental sources or other bacterial species, indicating that raw foods in the local market may serve as a reservoir of resistance-encoding genetic elements from which such elements are disseminated to various microbial pathogens. These findings suggest a need to perform periodic but comprehensive surveillance of multidrug-resistant bacterial pathogens and the major antimicrobial resistance genes in common food products, so as to disrupt the transmission routes of such organisms and the resistance-encoding genetic elements that they harbour.

Hong Kong↗

Harnessing the mouse to unravel the genetics of human disease.

Complex traits, i.e. those with multiple genetic and environmental determinants, represent the greatest challenge for genetic analysis, largely due to the difficulty of isolating the effects of any one gene amid the noise of other genetic and environmental influences. Methods exist for detecting and mapping the Quantitative Trait Loci (QTLs) that influence complex traits. However, once mapped, gene identification commonly involves reduction of focus to single candidate genes or isolated chromosomal regions. To reach the next level in unraveling the genetics of human disease will require moving beyond the focus on one gene at a time, to explorations of pleiotropism, epistasis and environment-dependency of genetic effects. Genetic interactions and unique environmental features must be as carefully scrutinized as are single gene effects. No one genetic approach is likely to possess all the necessary features for comprehensive analysis of a complex disease. Rather, the entire arsenal of behavioral genomic and other approaches will be needed, such as random mutagenesis, QTL analyses, transgenic and knockout models, viral mediated gene transfer, pharmacological analyses, gene expression assays, antisense approaches and importantly, revitalization of classical genetic methods. In our view, classical breeding designs are currently underutilized, and will shorten the distance to the target of understanding the complex genetic and environmental interactions associated with disease. We assert that unique combinations of classical approaches with current behavioral and molecular genomic approaches will more rapidly advance the field.

Animals↗

Genetic diversity among Mycoplasma species bovine group 7: clonal isolates from an outbreak of polyarthritis, mastitis, and abortion in dairy cattle.

A comprehensive genetic analysis of 60 Mycoplasma sp. bovine group 7 isolates from different geographic origins and epidemiological settings is presented. Twenty-four isolates were recovered from the joints of calves during sporadic episodes of polyarthritis in geographically distinct regions of Queensland and New South Wales, Australia, including two clones of the type strain PG5O. A further three Australian isolates were also recovered from the tympanic bulla, retropharyngeal lymph node and the lung and another three isolates had unconfirmed histories. Six isolates originated from Germany, Portugal, Nigeria, and France. Twenty-four epidemiologically related isolates of Mycoplasma sp. bovine group 7 were recovered from multiple tissue sites and body fluids of infected calves with polyarthritis, mastitic milk, and from the stomach contents, lung and liver from aborted foetuses in three large, centrally managed dairy herds in New South Wales, Australia. Restriction endonuclease analysis (REA) of genomic DNA differentiated 29 Cfol profiles among these 60 isolates and grouped all 24 epidemiologically related isolates in a defined pattern showing a clonal origin. Three isolates of this clonal cluster were recovered from mastitic milk and the synovial exudate of clinically-affected calves and appeared sporadically for periods up to 18 months after the initial outbreak of polyarthritis indicating a persistent, close association of the organism with cattle in these herds. The Cfol profile representative of the clonal cluster was distinguishable from profiles of isolates recovered from multiple, unrelated cases of polyarthritis in Queensland and New South Wales and from other countries. All 24 isolates from the clonal cluster possessed a plasmid (pBG7AU) with a molecular size of 1022 bp. DNA sequence analysis of pBG7AU identified two open reading frames sharing 81 and 99% DNA sequence similarity with hypothetical replication control proteins A and B respectively, previously described in plasmid pADB201 isolated from M. mycoides subspecies mycoides. Other isolates of bovine group 7, epidemiologically unrelated to the clonal cluster, including two clones of the type strain PG5O, possessed a similar-sized plasmid. These data confirm that Mycoplasma sp. bovine group 7 is capable of migrating to, and multiplying within, different tissue sites within a single animal and among different animals within a herd.

Abortion, Veterinary↗