Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Signature”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

An oxidative stress - and immunotherapy-related six-gene signature defines immune subtypes and predicts prognosis and immunotherapy response in hepatocellular carcinoma.

BACKGROUND: Oxidative stress and the tumor immune microenvironment jointly shape hepatocellular carcinoma (HCC) progression and response to immunotherapy, yet integrated biomarkers linking these processes are lacking. METHODS: Transcriptomic and clinical data from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) datasets were used to identify oxidative stress- and immunotherapyrelated differentially expressed genes (OSIRDEGs). Functional enrichment, weighted gene co-expression network analysis (WGCNA) and LASSO-Cox regression were used to construct a prognostic signature. Consensus clustering, TIDE, CIBERSORT and ssGSEA characterized immune phenotypes. Somatic mutation, copy-number and drug-response data were integrated to assess genomic alterations and drug sensitivity. Expression of model genes was validated by qRT-PCR and western blotting in HCC cell lines. RESULTS: We identified 24 OSIRDEGs enriched in cell-cycle and mitotic pathways. WGCNA intersection yielded 18 module genes, from which a six-gene signature (BUB1B, CDKN2A, CENPE, HMMR, PTTG1, SPP1) was derived. The signature robustly stratified patients into high- and low-risk groups with significantly different progression-free and disease-free survival in both TCGA-LIHC and GSE14520. Based on signature expression, two molecular subtypes were defined, exhibiting distinct survival, immune landscapes and predicted immunotherapy responsiveness. Model genes harbored recurrent alterations and showed significant correlations with anticancer agents. All six genes were upregulated at mRNA and protein levels in metastatic HCC cell lines versus normal hepatocytes. CONCLUSIONS: We systematically explored the landscape of OSIRDEGs in HCC, and proposed a validated six-gene signature that refines prognostic stratification, delineates immunerelevant HCC subtypes and highlights candidate biomarkers for therapeutic selection and mechanistic investigation.

Humans↗

Construction of molecular signatures based on the co-expression network of NECSO-related gene TRPM4 and its prognostic value in hepatocellular carcinoma.

BACKGROUND: Hepatocellular carcinoma (HCC) demonstrates significant prognostic variability that is not entirely accounted for by traditional staging systems. Necrosis by sodium overload (NECSO) is an emerging programmed cell death pathway, but its clinical relevance in HCC remains undefined. Therefore, this study aimed to identify TRPM4-associated core genes, develop and validate a prognostic signature, and investigate its relationship with the tumor immune microenvironment, tumor mutational burden, and single-cell expression patterns in HCC. METHODS: We integrated transcriptomic, clinical, and mutational datasets from The Cancer Genome Atlas-Liver Hepatocellular Carcinoma (TCGA-LIHC) (n=421) and Gene Expression Omnibus (GEO) cohorts (n=115) to identify genes co-expressed with TRPM4-a key NECSO mediator-and those differentially expressed in HCC. A prognostic signature was developed using least absolute shrinkage and selection operator (LASSO)-Cox regression and validated through survival analysis, time-dependent receiver operating characteristic (ROC) curves, and multivariate Cox regression analysis. The immune landscape was characterized using CIBERSORT, somatic mutation data were used to calculate tumor mutational burden (TMB) and assess its correlation with the risk score, and single-cell RNA sequencing (scRNA-seq) resolved cell-type-specific expression patterns. RESULTS: From 294 TRPM4-associated core genes, we identified an 11-gene signature (BRSK1, MMP1, GRIN2D, GP6, MYOM2, N4BP3, CCDC112, TSEN54, MAP3K9, SPP1, B3GNT4) that independently predicted overall survival (OS) (hazard ratio =5.419, P<0.001) with areas under the curve (AUCs) of 0.779, 0.693, and 0.701 at 1, 3, and 5 years. These values were superior or comparable to conventional clinicopathologic variables after direct comparison. High-risk patients exhibited an immunosuppressive microenvironment, characterized by enrichment of M0 macrophage, a higher M2/M1 ratio (P<0.001) and distinct immune checkpoint profiles. When integrated with TMB, the prognostic stratification was further refined: high-TMB/high-risk patients had poorest outcomes (median OS, 15.3 months), while low-TMB/low-risk patients had the most favorable survival (median OS, 68.7 months). Single-cell analysis revealed that MMP1 was induced in cancer-associated fibroblasts (CAFs) and SPP1 was downregulated in macrophages, single-cell risk scores confirmed TAFs and macrophages as the main contributors to the prognostic model. CONCLUSIONS: The TRPM4-centered 11-gene signature provides robust and independent prognostic stratification in HCC by integrating immune, mutational, and single-cell features. This signature serves as a potential tool for prognostic evaluation and may help inform immunotherapeutic strategies for HCC.

Hepatocellular carcinoma (HCC)↗

The molecular signature of selection underlying human adaptations.

In the last decade, advances in human population genetics and comparative genomics have resulted in important contributions to our understanding of human genetic diversity and genetic adaptation. For the first time, we are able to reliably detect the signature of natural selection from patterns of DNA polymorphism. Identifying the effects of natural selection in this way provides a crucial piece of evidence needed to support hypotheses of human adaptation. This review provides a detailed description of the theory and analytical approaches used to detect signatures of natural selection in the human genome. We discuss these methods in relation to four classic human traits--skin color, the Duffy blood group, bitter-taste sensation, and lactase persistence. By highlighting these four traits we are able to discuss the ways in which analyses of DNA polymorphism can lead to inferences regarding past histories of selection. Specifically, we can infer the importance of specific regimes of selection (i.e. directional selection, balancing selection, and purifying selection) in the evolution of a trait because these different types of selection leave different patterns of DNA polymorphism. In addition, we demonstrate how these types of data can be used to estimate the time frame in which selection operated on a trait. As the field has advanced, a general issue that has come to the forefront is how specific demographic events in human history, such as population expansions, bottlenecks, and subdivision of populations, have also left a signature across the genome that can interfere with our detection of the footprint of selection at particular genes. Therefore, we discuss this general problem with respect to the four traits reviewed here, and describe the ways in which the signature of selection can be teased from a background signature of demographic history. Finally, we move from a discussion of analyses of selection motivated by a "candidate-gene" approach, in which a priori information led to the analysis of specific gene, to discussion of "genome-scanning" approaches that are directed at discovering new genes that have been under positive selection. Such scans can be designed to detect those genes that have been positively selected in our divergence from chimpanzees, as well as those genes that have been under selection as human populations have migrated, differentiated, and adapted to specific geographic environments. We predict that both approaches will be applied in the future, enabling a greater insight into human species-wide adaptations, as well as the specific adaptations of human populations.

Animals↗

Cancer-associated molecular signature in the tissue samples of patients with cirrhosis.

Several types of aggressive cancers, including hepatocellular carcinoma (HCC), often arise as a multifocal primary tumor. This suggests a high rate of premalignant changes in noncancerous tissue before the formation of a solitary tumor. Examination of the messenger RNA expression profiles of tissue samples derived from patients with cirrhosis of various etiologies by complementary DNA (cDNA) microarray indicated that they can be grossly separated into two main groups. One group included hepatitis B and C virus infections, hemochromatosis, and Wilson's disease. The other group contained mainly alcoholic liver disease, autoimmune hepatitis, and primary biliary cirrhosis. Analysis of these two groups by the cross-validated leave-one-out machine-learning algorithms revealed a molecular signature containing 556 discriminative genes (P <.001). It is noteworthy that 273 genes in this signature (49%) were also significantly altered in HCC (P <.001). Many genes were previously known to be related to HCC. The 273-gene signature was validated as cancer-associated genes by matching this set to additional independent tumor tissue samples from 163 patients with HCC, 56 patients with lung carcinoma, and 38 patients with breast carcinoma. From this signature, 30 genes were altered most significantly in tissue samples from high-risk individuals with cirrhosis and from patients with HCC. Among them, 12 genes encoded secretory proteins found in sera. In conclusion, we identified a unique gene signature in the tissue samples of patients with cirrhosis, which may be used as candidate markers for diagnosing the early onset of HCC in high-risk populations and may guide new strategies for chemoprevention. Supplementary material for this article can be found on the HEPATOLOGY website (http://interscience.wiley.com/jpages/0270-9139/suppmat/index.html).

Antigens, Neoplasm↗

Effects of elemental composition on the incorporation of dietary nitrogen and carbon isotopic signatures in an omnivorous songbird.

The use of stable isotopes to infer diet requires quantifying the relationship between diet and tissues and, in particular, knowing of how quickly isotopes turnover in different tissues and how isotopic concentrations of different food components change (discriminate) when incorporated into consumer tissues. We used feeding trials with wild-caught yellow-rumped warblers (Dendroica coronata) to determine delta15N and delta13C turnover rates for blood, delta15N and delta13C diet-tissue discrimination factors, and diet-tissue relationships for blood and feathers. After 3 weeks on a common diet, 36 warblers were assigned to one of four diets differing in the relative proportion of fruit and insects. Plasma half-life estimates ranged from 0.4 to 0.7 days for delta13C and from 0.5 to 1.7 days for delta15N . Half-life did not differ among diets. Whole blood half-life for delta13C ranged from 3.9 to 6.1 days. Yellow-rumped warbler tissues were enriched relative to diet by 1.7-3.6% for nitrogen isotopes and by -1.2 to 4.3% for carbon isotopes, depending on tissue and diet. Consistent with previous studies, feathers were the most enriched and whole blood and plasma were the least enriched or, in the case of carbon, slightly depleted relative to diet. In general, tissues were more enriched relative to diet for birds on diets with high percentages of insects. For all tissues, carbon and nitrogen isotope discrimination factors increased with carbon and nitrogen concentrations of diets. The isotopic signature of plasma increased linearly with the sum of the isotopic signature of the diet and the discrimination factor. Because the isotopic signature of tissues depends on both elemental concentration and isotopic signature of the diet, attempts to reconstruct diet from stable isotope signatures require use of mixing models that incorporate elemental concentration.

Animals↗

A method for simple identification of signature peptides derived from polyUb-K48 and K63 by MALDI-TOF MS and chemically assisted MS/MS fragmentation.

A simple method is described to identify signature peptides derived from polyubiquitin (polyUb) chains. The method is based on MALDI-TOF MS/MS analysis after chemically assisted fragmentation, and works on peptides isolated from polyacrylamide gels. PolyUb chains branched at K48 and K63 were chosen as models for Ub-protein conjugates. They were resolved by SDS-PAGE, and their tryptic peptides (in-gel-trypsinolysis) derivatized with 3-sulfopropinic acid NHSester to obtain chemically assisted fragmentation during the MS/MS analysis. PolyUb-K63 produced a single peptide identified as (55)TLSDYNIQK(63) (GG)ESTLHLVLR(72). PolyUb-K48 produced two branched signature peptides identified as (43)LIFAGK(48)(GG)QLEDGR(54) and (43)LIFAGK(48)(LRGG)QLEDGR(54). The recovery of signature peptide with LRGG as branched chain underscores the need to take limited proteolysis into account in the search for detection of ubiquitinated peptides in proteomics studies. In conclusion, a simple method has been described allowing the identification of signature peptides, which are diagnostic markers of the majority of polyUb-conjugated proteins. In principle, the method should be applicable also for other more rare signature peptides.

Amino Acid Sequence↗

Signature function for predicting resonant and attenuant population 2-cycles.

Populations are either enhanced via resonant cycles or suppressed via attenuant cycles by periodic environments. We develop a signature function for predicting the response of discretely reproducing populations to 2-periodic fluctuations of both a characteristic of the environment (carrying capacity), and a characteristic of the population (inherent growth rate). Our signature function is the sign of a weighted sum of the relative strengths of the oscillations of the carrying capacity and the demographic characteristic. Periodic environments are deleterious for populations when the signature function is negative. However, positive signature functions signal favorable environments. We compute the signature functions of six classical discrete-time single species population models, and use the functions to determine regions in parameter space that are either favorable or detrimental to the populations. The two-parameter classical models include the Ricker, Beverton-Holt, Logistic, and Maynard Smith models.

Animals↗

Thyroid hormone deprivation creates an immunological signature in the mouse liver, involving Kupffer cell presentation as the mouse ages.

PURPOSE: Aging is associated with an increased prevalence of chronic liver diseases suggesting impaired immune and metabolic function. In addition, thyroid hormone (TH) impacts liver physiology and TH deprivation or excess negatively affect organ maintenance. However, whether age-dependent consequences of TH alterations are reflected in a liver-specific adaptation is unknown so far. The present study aimed to characterize the impact of TH deprivation or excess on the liver transcriptome during aging. METHODS: Five- and 21-month-old male C57BL/6 mice were exposed either to chronic TH deprivation or to chronic TH excess and compared to control treatment by microarray-based liver transcriptome analysis. RESULTS: Significant roles of both TH state and age became obvious: Bioinformatic analysis of the liver transcriptome data revealed an age-dependent immune signature by chronic TH deprivation, an age-dependent immune and metabolic signature independent of exogenous TH modulation, as well as an age-dependent metabolic signature by chronic TH excess. Published data of single cell transcriptomic atlas characterizing aging tissues in the mouse were compared with our data and revealed Kupffer cell presentation in the immunological signature by TH deprivation during aging. Literature data for four prominent differentially expressed genes, namely C1qb, C3ar1, Ctss, and Msr1, revealed that the complement system, extracellular matrix remodelling, as well as the proinflammatory phenotype of Kupffer cells are altered by TH deprivation during aging. CONCLUSION: In conclusion, our study illuminates the interplay between TH deprivation, aging, and liver transcriptome signatures, highlighting potential implications for immune function and tissue maintenance, particularly through the modulation of Kupffer cell presentation.

Animals↗

Electron paramagnetic resonance dynamic signatures of TAR RNA-small molecule complexes provide insight into RNA structure and recognition.

Electron paramagnetic resonance (EPR) spectroscopy was utilized to investigate the correlation between RNA structure and RNA internal dynamics in complexes of HIV-1 TAR RNA with small molecules. TAR RNAs containing single nitroxide spin-labels in the 2'-position of U23, U25, U38, or U40 were incubated with compounds known to inhibit TAR-Tat complex formation. The combined changes in nucleotide mobility at all four sites, as monitored by their EPR spectral width, yield a dynamic signature for each compound. The multicyclic dyes Hoechst 33258, DAPI, and berenil bind to TAR RNA in a similar manner and gave nearly identical signatures. Different signatures were obtained for the acridine derivative CGP 40336A and the aminoglycoside antibiotic neomycin, which bind to different regions of the RNA. The dynamic signature for guanidinoneomycin was remarkably similar to that obtained for argininamide and is evidence for guanidinoneomycin binding to the same site as arginine 52 of the Tat protein, rather than to the neomycin binding site. The data presented here show that the dynamic signatures provide strong insights into RNA structure and recognition and demonstrate the value of EPR spectroscopy for the investigation of small molecule binding to RNA.

Acridines↗

Shape signatures: a new approach to computer-aided ligand- and receptor-based drug design.

A unifying principle of rational drug design is the use of either shape similarity or complementarity to identify compounds expected to be active against a given target. Shape similarity is the underlying foundation of ligand-based methods, which seek compounds with structure similar to known actives, while shape complementarity is the basis of most receptor-based design, where the goal is to identify compounds complementary in shape to a given receptor. These approaches can be extended to include molecular descriptors in addition to shape, such as lipophilicity or electrostatic potential. Here we introduce a new technique, which we call shape signatures, for describing the shape of ligand molecules and of receptor sites. The method uses a technique akin to ray-tracing to explore the volume enclosed by a ligand molecule, or the volume exterior to the active site of a protein. Probability distributions are derived from the ray-trace, and can be based solely on the geometry of the reflecting ray, or may include joint dependence on properties, such as the molecular electrostatic potential, computed over the surface. Our shape signatures are just these probability distributions, stored as histograms. They converge rapidly with the length of the ray-trace, are independent of molecular orientation, and can be compared quickly using simple metrics. Shape signatures can be used to test for both shape similarity between compounds and for shape complementarity between compounds and receptors and thus can be applied to problems in both ligand- and receptor-based molecular design. We present results for comparisons between small molecules of biological interest and the NCI Database using shape signatures under two different metrics. Our results show that the method can reliably extract compounds of shape (and polarity) similar to the query molecules. We also present initial results for a receptor-based strategy using shape signatures, with application to the design of new inhibitors predicted to be active against HIV protease.

Binding Sites↗

Identification and validation of a previously missed mutational signature in colorectal cancer.

Mutational signature analysis has enhanced our understanding of mutagenic processes. In a recent study, we analyzed 802 microsatellite-stable colorectal cancers (CRC) and identified a de novo signature, SBS_D, which was decomposed into SBS18. Here, we re-evaluate this decomposition and provide evidence that SBS_D represents a distinct mutational process from SBS18. Through an analysis of 2,616 CRCs across three independent cohorts, we demonstrate that SBS_D is consistently present, suggesting this signature may have been previously overlooked. We illustrate that the pattern of SBS_D better aligns with signatures associated with deficiencies in DNA repair, despite evidence that SBS_D is not driven by canonical defects in these DNA repair pathways. Overall, this study identifies a previously unrecognized mutational signature in DNA repair-proficient CRC and proposes that its etiology may be linked to DNA repair infidelity emerging late in tumor development. SBS_D has been submitted to the COSMIC database and provisionally designated as SBS111.

Colorectal Neoplasms↗

Machine learning-based analysis of oral rinse samples to identify candidate proteomic signatures for severe periodontitis: a pilot study.

This pilot study investigated whether candidate protein signatures from oral rinse samples can distinguish patients with severe periodontitis (stage III/IV) and its subtypes, generalized and localized periodontitis, from non-periodontitis controls. Participants rinsed with phosphate-buffered saline, and samples were analyzed using a Proximity Extension Assay targeting 92 inflammatory and 92 immuno-oncology proteins. A machine learning approach using repeated nested cross-validation and SHAP was implemented to identify protein signatures. The study included 38 patients (18 with localized periodontitis and 20 with generalized periodontitis) and 16 controls. After data preprocessing, 54 samples and 141 proteins were retained. Proteins Gal-1, HGF, TNFSF14, CD27, and ARG1 distinguished periodontitis from controls (ROC-AUC&#x2009;=&#x2009;0.85, 95% CI 0.82, 0.87). For generalized periodontitis, we found a protein signature including TNFSF14, Gal-1, STAMBP, MUC-16, S100A12, HGF, CASP-8, CD27, LAP TGF-&#x3b2;1, TNFRSF9, and uPA (ROC-AUC&#x2009;=&#x2009;0.92, 95% CI 0.90, 0.94). For localized periodontitis, we identified ARG1 (ROC-AUC&#x2009;=&#x2009;0.72, 95% CI 0.68, 0.76). No proteomic signature distinguishing generalized periodontitis from localized periodontitis was identified. This pilot study indicated that oral rinses are suitable for proteomic profiling, and there was a putative protein signature that could differentiate periodontitis, generalized periodontitis, and localized periodontitis from controls. These findings warrant validation in larger independent cohorts, including a clearly defined gingivitis group, before real-world non-invasive screening applications can be considered.

Humans↗

Signature whistle shape conveys identity information to bottlenose dolphins.

Bottlenose dolphins (Tursiops truncatus) develop individually distinctive signature whistles that they use to maintain group cohesion. Unlike the development of identification signals in most other species, signature whistle development is strongly influenced by vocal learning. This learning ability is maintained throughout life, and dolphins frequently copy each other's whistles in the wild. It has been hypothesized that signature whistles can be used as referential signals among conspecifics, because captive bottlenose dolphins can be trained to use novel, learned signals to label objects. For this labeling to occur, signature whistles would have to convey identity information independent of the caller's voice features. However, experimental proof for this hypothesis has been lacking. This study demonstrates that bottlenose dolphins extract identity information from signature whistles even after all voice features have been removed from the signal. Thus, dolphins are the only animals other than humans that have been shown to transmit identity information independent of the caller's voice or location.

Animal Communication↗

Clarifying the catalytic roles of conserved residues in the amidase signature family.

Fatty acid amide hydrolase (FAAH) is a mammalian integral membrane enzyme responsible for the hydrolysis of a number of neuromodulatory fatty acid amides, including the endogenous cannabinoid anandamide and the sleep-inducing lipid oleamide. FAAH belongs to a large class of hydrolytic enzymes termed the "amidase signature family," whose members are defined by a conserved stretch of approximately 130 amino acids termed the "amidase signature sequence." Recently, site-directed mutagenesis studies of FAAH have targeted a limited number of conserved residues in the amidase signature sequence of the enzyme, identifying Ser-241 as the catalytic nucleophile and Lys-142 as an acid/base catalyst. The roles of several other conserved residues with potentially important and/or overlapping catalytic functions have not yet been examined. In this study, we have mutated all potentially catalytic residues in FAAH that are conserved among members of the amidase signature family, and have assessed their individual roles in catalysis through chemical labeling and kinetic methods. Several of these residues appear to serve primarily structural roles, as their mutation produced FAAH variants with considerable catalytic activity but reduced expression in prokaryotic and/or eukaryotic systems. In contrast, five mutations, K142A, S217A, S218A, S241A, and R243A, decreased the amidase activity of FAAH greater than 100-fold without detectably impacting the structural integrity of the enzyme. The pH rate profiles, amide/ester selectivities, and fluorophosphonate reactivities of these mutants revealed distinct catalytic roles for each residue. Of particular interest, one mutant, R243A, displayed uncompromised esterase activity but severely reduced amidase activity, indicating that the amidase and esterase efficiencies of FAAH can be functionally uncoupled. Collectively, these studies provide evidence that amidase signature enzymes represent a large class of serine-lysine catalytic dyad hydrolases whose evolutionary distribution rivals that of the catalytic triad superfamily.

Amidohydrolases↗

Identification of a prognostic signature consisting of three macrophage-related genes for glioblastoma based on bulk and single-cell transcriptomes analyses.

BACKGROUND: Tumor-associated macrophages have been implicated in the progression and treatment resistance of glioblastoma (GBM). This study aimed to identify macrophage-related genes associated with prognosis and therapeutic response in GBM. MATERIALS AND METHODS: Bulk RNA-seq data from 533 patients with GBM were downloaded from the Cancer Genome Atlas (TCGA) and Chinese Glioma Genome Atlas (CGGA) databases. Bioinformatic tools were used to detect the co-expression gene modules associated with the infiltration of immune cells, identify a prognostic macrophage-related gene signature, and explore their association with sensitivity to chemotherapeutic drugs and immune checkpoint blockade. Single-cell RNA-seq data and multiplexed immunofluorescence were used to validate ISG20 expression (a member of the identified gene signature) in macrophages. RESULTS: We detected gene modules associated with macrophages and identified a signature consisting of three macrophage-related genes (ISG20, PARP12 and IFIT5) in the discovery set (TCGA-GBM, n&#x2009;=&#x2009;159), and validated its prognostic value in the validation set (CGGA-GBM, n&#x2009;=&#x2009;374). This gene signature demonstrated favorable accuracy in predicting prognosis and resistance of immuno- and chemo-therapy. The co-expression of ISG20 and PD-1 in macrophages was verified by single-cell RNA-seq data and multiplex immunofluorescence. CONCLUSIONS: This study presents a macrophage-related gene signature to predict prognosis and therapeutic response in GBM. ISG20, PARP12 and IFIT5 are interferon-stimulated genes, and further investigations may provide new insights into the interplay between macrophages and interferon signaling in GBM.

Humans↗

The phylogeny and signature sequences characteristics of Fibrobacteres, Chlorobi, and Bacteroidetes.

Fibrobacteres, Chlorobi, and Bacteroidetes (FCB group) comprise three main bacterial phyla recognized on the basis of 16S rRNA trees. Presently, there are no distinctive biochemical or molecular characteristics known that can distinguish these bacteria from other bacterial phyla. The relationship of these bacteria to other phyla is also not known. This review describes many signatures, consisting of defined and conserved inserts in widely distributed proteins, that provide distinctive molecular markers for these groups of bacteria. These signatures serve to clarify the evolutionary relationship between members of the FCB group, and to other bacterial phyla. A 4 aa insert in DNA Gyrase B (GyrB) and a 45 aa insert in the SecA proteins are uniquely shared by various Bacteroidetes species. The insert in GyrB is present in all Bacteroidetes species (>100) covering different orders and families, indicating that it is a distinctive characteristic of the group. Three signatures consisting of an 18 aa insert in ATPase alpha-subunit, an 8-9 aa insert in the FtsK protein and a 1 aa insert in the UvrB protein are commonly shared only by the Bacteroidetes and Chlorobi homologs providing evidence that these two groups are specifically related to each other. Two additional inserts in the RNA polymerase beta'-subunit (5-7 aa) and Serine hydroxymethyl-transferase (14-16 aa), which are commonly present in various Bacteroidetes, Chlorobi, and Fibrobacteres homologs, but not any other bacteria, provide evidence that these groups shared a common ancestor exclusive of all other bacteria. The FCB groups of bacteria are indicated to have diverged from this common ancestor in the following order: Fibrobacteres --> Chlorobi --> Bacteriodetes. The inferences from signature sequences are strongly supported by phylogenetic analyses. These observations suggest that the FCB groups of bacteria should be placed in a single phylum rather than three distinct phyla. Signature sequences in a number of other proteins provide evidence that the FCB group of bacteria diverged at a similar time as the Chlamydiae group, and that the Spirochetes and Aquificales groups are its closest relatives.

Adenosine Triphosphatases↗

Signature pattern analysis: a method for assessing viral sequence relatedness.

Signature pattern analysis identifies particular sites in amino acid or nucleic acid alignments of variable sequences that are distinctly representative of a query set of sequences relative to a background set. We explore the merits of using signature patterns for analysis of HIV-1 (human immunodeficiency virus type 1) sequences in cases of epidemiological linkage and potential superinfection. For these purposes, query sets are viral sequences that are all derived from one HIV-1 infected individual, hence the signature pattern is the array of sites that are characteristic of the range of viral variants obtained from that person. Once a signature pattern has been objectively defined, it can be used to examine other viral sequences from other individuals for evidence of genetic relatedness. A computer program to facilitate this analysis, VESPA, is described and applied to sequence data gathered during the investigation of HIV-1 transmission in a dental practice. The implications of signature polymorphisms seen within an infected individual, and shared polymorphisms between linked individuals, are also considered. VESPA may also be applied to the molecular analysis of biological phenotypes.

Amino Acid Sequence↗

Reliable gene signatures for microarray classification: assessment of stability and performance.

MOTIVATION: Two important questions for the analysis of gene expression measurements from different sample classes are (1) how to classify samples and (2) how to identify meaningful gene signatures (ranked gene lists) exhibiting the differences between classes and sample subsets. Solutions to both questions have immediate biological and biomedical applications. To achieve optimal classification performance, a suitable combination of classifier and gene selection method needs to be specifically selected for a given dataset. The selected gene signatures can be unstable and the resulting classification accuracy unreliable, particularly when considering different subsets of samples. Both unstable gene signatures and overestimated classification accuracy can impair biological conclusions. METHODS: We address these two issues by repeatedly evaluating the classification performance of all models, i.e. pairwise combinations of various gene selection and classification methods, for random subsets of arrays (sampling). A model score is used to select the most appropriate model for the given dataset. Consensus gene signatures are constructed by extracting those genes frequently selected over many samplings. Sampling additionally permits measurement of the stability of the classification performance for each model, which serves as a measure of model reliability. RESULTS: We analyzed a large gene expression dataset with 78 measurements of four different cartilage sample classes. Classifiers trained on subsets of measurements frequently produce models with highly variable performance. Our approach provides reliable classification performance estimates via sampling. In addition to reliable classification performance, we determined stable consensus signatures (i.e. gene lists) for sample classes. Manual literature screening showed that these genes are highly relevant to our gene expression experiment with osteoarthritic cartilage. We compared our approach to others based on a publicly available dataset on breast cancer. AVAILABILITY: R package at http://www.bio.ifi.lmu.de/~davis/edaprakt

Algorithms↗