Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

Gene expression signature of estrogen receptor alpha status in breast cancer.

BACKGROUND: Estrogens are known to regulate the proliferation of breast cancer cells and to modify their phenotypic properties. Identification of estrogen-regulated genes in human breast tumors is an essential step toward understanding the molecular mechanisms of estrogen action in cancer. To this end we generated and compared the Serial Analysis of Gene Expression (SAGE) profiles of 26 human breast carcinomas based on their estrogen receptor alpha (ER) status. Thus, producing a breast cancer SAGE database of almost 2.5 million tags, representing over 50,000 transcripts. RESULTS: We identified 520 transcripts differentially expressed between ERalpha-positive (+) and ERalpha-negative (-) primary breast tumors (Fold change >or= 2; p < 0.05). Furthermore, we identified 220 high-affinity Estrogen Responsive Elements (EREs) distributed on the promoter regions of 163 out of the 473 up-modulated genes in ERalpha (+) breast tumors. In brief, we observed predominantly up-regulation of cell growth related genes, DNA binding and transcription factor activity related genes based on Gene Ontology (GO) biological functional annotation. GO terms over-representation analysis showed a statistically significant enrichment of various transcript families including: metal ion binding related transcripts (p = 0.011), calcium ion binding related transcripts (p = 0.033) and steroid hormone receptor activity related transcripts (p = 0.031). SAGE data associated with ERalpha status was compared with reported information from breast cancer DNA microarrays studies. A significant proportion of ERalpha associated gene expression changes was validated by this cross-platform comparison. However, our SAGE study also identified novel sets of genes as highly expressed in ERalpha (+) invasive breast tumors not previously reported. These observations were further validated in an independent set of human breast tumors by means of real time RT-PCR. CONCLUSION: The integration of the breast cancer comparative transcriptome analysis based on ERalpha status coupled to the genome-wide identification of high-affinity EREs and GO over-representation analysis, provide useful information for validation and discovery of signaling networks related to estrogen response in this malignancy.

Biomarkers, Tumor↗

Comparative mapping of expressed sequence tags containing microsatellites in rainbow trout (Oncorhynchus mykiss).

BACKGROUND: Comparative genomics, through the integration of genetic maps from species of interest with whole genome sequences of other species, will facilitate the identification of genes affecting phenotypes of interest. The development of microsatellite markers from expressed sequence tags will serve to increase marker densities on current salmonid genetic maps and initiate in silico comparative maps with species whose genomes have been fully sequenced. RESULTS: Eighty-nine polymorphic microsatellite markers were generated for rainbow trout of which at least 74 amplify in other salmonids. Fifty-five have been associated with functional annotation and 30 were mapped on existing genetic maps. Homologous sequences were identified for 20 of the EST containing microsatellites to identify comparative assignments within the tetraodon, mouse, and/or human genomes. CONCLUSION: The addition of microsatellite markers constructed from expressed sequence tag data will facilitate the development of high-density genetic maps for rainbow trout and comparative maps with other salmonids and better studied species.

Alleles↗

Differences in the evolutionary history of disease genes affected by dominant or recessive mutations.

BACKGROUND: Global analyses of human disease genes by computational methods have yielded important advances in the understanding of human diseases. Generally these studies have treated the group of disease genes uniformly, thus ignoring the type of disease-causing mutations (dominant or recessive). In this report we present a comprehensive study of the evolutionary history of autosomal disease genes separated by mode of inheritance. RESULTS: We examine differences in protein and coding sequence conservation between dominant and recessive human disease genes. Our analysis shows that disease genes affected by dominant mutations are more conserved than those affected by recessive mutations. This could be a consequence of the fact that recessive mutations remain hidden from selection while heterozygous. Furthermore, we employ functional annotation analysis and investigations into disease severity to support this hypothesis. CONCLUSION: This study elucidates important differences between dominantly- and recessively-acting disease genes in terms of protein and DNA sequence conservation, paralogy and essentiality. We propose that the division of disease genes by mode of inheritance will enhance both understanding of the disease process and prediction of candidate disease genes in the future.

Animals↗

Inferring direct regulatory targets from expression and genome location analyses: a comparison of transcription factor deletion and overexpression.

BACKGROUND: Effects on gene expression due to environmental or genetic changes can be easily measured using microarrays. However, indirect effects on expression can be substantial. The indirect effects of a perturbation need to be distinguished from the direct effects if we are to understand the structure and behavior of regulatory networks. RESULTS: The most direct way to perturb a transcriptional network is to alter transcription factor activity. Here, for the first time, we compare expression changes and genomic binding in a simple regulon under conditions of both low and high transcription factor activity. Specifically, we assessed the effects on expression and binding due to deletion of the yeast LEU3 transcription factor gene and effects due to elevation of Leu3 activity. Leu3 activity was elevated through overexpression and the introduction of a mutation that renders the protein constitutively active. Genes that are bound and/or regulated by Leu3 under one or both conditions were characterized in terms of their functional annotations and their predicted potential to be bound by Leu3. We also assessed the evolutionary conservation of the predicted binding potential using a novel alignment-independent method. Both perturbations yield genes that are likely to be direct targets of Leu3, including most of the classically defined targets. Additional direct targets are identified by each of the methods. However, experimental and computational criteria suggest that most genes whose expression is affected by the Leu3 genotype are unlikely to be regulated by binding of the protein. CONCLUSION: Most genes that are differentially expressed by Leu3 are not direct targets despite the exceptional simplicity of the regulon, and the unusually direct nature of the perturbations investigated. These conclusions are reached through computational analyses that support and extend chromatin immunoprecipitation data on the identities of direct targets. These results have implications for the interpretation of expression experiments, especially in cases for which chromatin immunoprecipitation data are unavailable, incomplete, or ambiguous.

2-Isopropylmalate Synthase↗

Development of a chicken 5 K microarray targeted towards immune function.

BACKGROUND: The development of microarray resources for the chicken is an important step in being able to profile gene expression changes occurring in birds in response to different challenges and stimuli. The creation of an immune-related array is highly valuable in determining the host immune response in relation to infection with a wide variety of bacterial and viral diseases. RESULTS: Here we report the development of chicken immune-related cDNA libraries and the subsequent construction of a microarray containing 5190 elements (in duplicate). Clones on the array originate from tissues known to contain high levels of cells related to the immune system, namely Bursa, Peyers patch, thymus and spleen. Represented on the array are genes that are known to cluster with existing chicken ESTs as well as genes that are unique to our libraries. Some of these genes have no known homologies and represent novel genes in the chicken collection. A series of reference genes (ie. genes of known immune function) are also present on the array. Functional annotation data is also provided for as many of the genes on the array as is possible. CONCLUSION: Six new chicken immune cDNA libraries have been created and nearly 10,000 sequences submitted to GenBank [GenBank: AM063043-AM071350; AM071520-AM072286; AM075249-AM075607]. A 5 K immune-related array has been developed from these libraries. Individual clones and arrays are available from the ARK-Genomics resource centre.

Animals↗

Development of ESTs from chickpea roots and their use in diversity analysis of the Cicer genus.

BACKGROUND: Chickpea is a major crop in many drier regions of the world where it is an important protein-rich food and an increasingly valuable traded commodity. The wild annual Cicer species are known to possess unique sources of resistance to pests and diseases, and tolerance to environmental stresses. However, there has been limited utilization of these wild species by chickpea breeding programs due to interspecific crossing barriers and deleterious linkage drag. Molecular genetic diversity analysis may help predict which accessions are most likely to produce fertile progeny when crossed with chickpea cultivars. While, trait-markers may provide an effective tool for breaking linkage drag. Although SSR markers are the assay of choice for marker-assisted selection of specific traits in conventional breeding populations, they may not provide reliable estimates of interspecific diversity, and may lose selective power in backcross programs based on interspecific introgressions. Thus, we have pursued the development of gene-based markers to resolve these problems and to provide candidate gene markers for QTL mapping of important agronomic traits. RESULTS: An EST library was constructed after subtractive suppressive hybridization (SSH) of root tissue from two very closely related chickpea genotypes (Cicer arietinum). A total of 106 EST-based markers were designed from 477 sequences with functional annotations and these were tested on C. arietinum. Forty-four EST markers were polymorphic when screened across nine Cicer species (including the cultigen). Parsimony and PCoA analysis of the resultant EST-marker dataset indicated that most accessions cluster in accordance with the previously defined classification of primary (C. arietinum, C. echinospermum and C. reticulatum), secondary (C. pinnatifidum, C. bijugum and C. judaicum), and tertiary (C. yamashitae, C. chrossanicum and C. cuneatum) gene-pools. A large proportion of EST alleles (45%) were only present in one or two of the accessions tested whilst the others were represented in up to twelve of the accessions tested. CONCLUSION: Gene-based markers have proven to be effective tools for diversity analysis in Cicer and EST diversity analysis may be useful in identifying promising candidates for interspecific hybridization programs. The EST markers generated in this study have detected high levels of polymorphism amongst both common and rare alleles. This suggests that they would be useful for allele-mining of germplasm collections for identification of candidate accessions in the search for new sources of resistance to pests / diseases, and tolerance to abiotic stresses.

Biomarkers↗

Establishment of the epithelial-specific transcriptome of normal and malignant human breast cells based on MPSS and array expression data.

INTRODUCTION: Diverse microarray and sequencing technologies have been widely used to characterise the molecular changes in malignant epithelial cells in breast cancers. Such gene expression studies to identify markers and targets in tumour cells are, however, compromised by the cellular heterogeneity of solid breast tumours and by the lack of appropriate counterparts representing normal breast epithelial cells. METHODS: Malignant neoplastic epithelial cells from primary breast cancers and luminal and myoepithelial cells isolated from normal human breast tissue were isolated by immunomagnetic separation methods. Pools of RNA from highly enriched preparations of these cell types were subjected to expression profiling using massively parallel signature sequencing (MPSS) and four different genome wide microarray platforms. Functional related transcripts of the differential tumour epithelial transcriptome were used for gene set enrichment analysis to identify enrichment of luminal and myoepithelial type genes. Clinical pathological validation of a small number of genes was performed on tissue microarrays. RESULTS: MPSS identified 6,553 differentially expressed genes between the pool of normal luminal cells and that of primary tumours substantially enriched for epithelial cells, of which 98% were represented and 60% were confirmed by microarray profiling. Significant expression level changes between these two samples detected only by microarray technology were shown by 4,149 transcripts, resulting in a combined differential tumour epithelial transcriptome of 8,051 genes. Microarray gene signatures identified a comprehensive list of 907 and 955 transcripts whose expression differed between luminal epithelial cells and myoepithelial cells, respectively. Functional annotation and gene set enrichment analysis highlighted a group of genes related to skeletal development that were associated with the myoepithelial/basal cells and upregulated in the tumour sample. One of the most highly overexpressed genes in this category, that encoding periostin, was analysed immunohistochemically on breast cancer tissue microarrays and its expression in neoplastic cells correlated with poor outcome in a cohort of poor prognosis estrogen receptor-positive tumours. CONCLUSION: Using highly enriched cell populations in combination with multiplatform gene expression profiling studies, a comprehensive analysis of molecular changes between the normal and malignant breast tissue was established. This study provides a basis for the identification of novel and potentially important targets for diagnosis, prognosis and therapy in breast cancer.

Biomarkers, Tumor↗

Prediction of unidentified human genes on the basis of sequence similarity to novel cDNAs from cynomolgus monkey brain.

BACKGROUND: The complete assignment of the protein-coding regions of the human genome is a major challenge for genome biology today. We have already isolated many hitherto unknown full-length cDNAs as orthologs of unidentified human genes from cDNA libraries of the cynomolgus monkey (Macaca fascicularis) brain (parietal lobe and cerebellum). In this study, we used cDNA libraries of three other parts of the brain (frontal lobe, temporal lobe and medulla oblongata) to isolate novel full-length cDNAs. RESULTS: The entire sequences of novel cDNAs of the cynomolgus monkey were determined, and the orthologous human cDNA sequences were predicted from the human genome sequence. We predicted 29 novel human genes with putative coding regions sharing an open reading frame with the cynomolgus monkey, and we confirmed the expression of 21 pairs of genes by the reverse transcription-coupled polymerase chain reaction method. The hypothetical proteins were also functionally annotated by computer analysis. CONCLUSIONS: The 29 new genes had not been discovered in recent explorations for novel genes in humans, and the ab initio method failed to predict all exons. Thus, monkey cDNA is a valuable resource for the preparation of a complete human gene catalog, which will facilitate post-genomic studies.

Animals↗

Expression profiling of the schizont and trophozoite stages of Plasmodium falciparum with a long-oligonucleotide microarray.

BACKGROUND: The worldwide persistence of drug-resistant Plasmodium falciparum, the most lethal variety of human malaria, is a global health concern. The P. falciparum sequencing project has brought new opportunities for identifying molecular targets for antimalarial drug and vaccine development. RESULTS: We developed a software package, ArrayOligoSelector, to design an open reading frame (ORF)-specific DNA microarray using the publicly available P. falciparum genome sequence. Each gene was represented by one or more long 70 mer oligonucleotides selected on the basis of uniqueness within the genome, exclusion of low-complexity sequence, balanced base composition and proximity to the 3' end. A first-generation microarray representing approximately 6,000 ORFs of the P. falciparum genome was constructed. Array performance was evaluated through the use of control oligonucleotide sets with increasing levels of introduced mutations, as well as traditional northern blotting. Using this array, we extensively characterized the gene-expression profile of the intraerythrocytic trophozoite and schizont stages of P. falciparum. The results revealed extensive transcriptional regulation of genes specialized for processes specific to these two stages. CONCLUSIONS: DNA microarrays based on long oligonucleotides are powerful tools for the functional annotation and exploration of the P. falciparum genome. Expression profiling of trophozoites and schizonts revealed genes associated with stage-specific processes and may serve as the basis for future drug targets and vaccine development.

Animals↗

Integration with the human genome of peptide sequences obtained by high-throughput mass spectrometry.

A crucial aim upon the completion of the human genome is the verification and functional annotation of all predicted genes and their protein products. Here we describe the mapping of peptides derived from accurate interpretations of protein tandem mass spectrometry (MS) data to eukaryotic genomes and the generation of an expandable resource for integration of data from many diverse proteomics experiments. Furthermore, we demonstrate that peptide identifications obtained from high-throughput proteomics can be integrated on a large scale with the human genome. This resource could serve as an expandable repository for MS-derived proteome information.

Amino Acid Sequence↗

Integration of GWAS and WGCNA reveals novel candidate genes for cottonseed oil content in Gossypium hirsutum L.

Genetic improvement of cottonseed oil content represents a crucial strategy for enhancing the comprehensive utilization of cotton. Here, genome-wide association study (GWAS) and weighted gene co-expression network analysis (WGCNA) were integrated to elucidate the genetic control underlying oil content. Phenotypic evaluation of 159 cotton accessions revealed extensive genetic variation, with kernel oil content ranging from 17.81% to 39.50%. Population structure analysis based on 20,213 single nucleotide polymorphisms (SNPs) classified the accessions into two major subpopulations. A total of 18 SNPs exhibited significant associations with oil content, two of which were stably detected across multiple environments using the FarmCPU model. Further haplotype analysis within linkage disequilibrium (LD) blocks confirmed a favorable haplotype on chromosome A05 that was strongly correlated with elevated oil content. Integration of publicly available transcriptome data from 11 ovule developmental stages with WGCNA identified modules significantly linked to oil content. Of the 74 candidate genes within LD intervals, 17 were assigned to WGCNA modules. Functional annotation and enrichment analyses highlighted four putative candidate genes (GH_A05G1503, GH_A05G1506, GH_A05G1531, and GH_A10G2150) involved in oil biosynthesis. These findings deepen our understanding of the genetic mechanisms governing cottonseed oil biosynthesis and lay a foundation for breeding high-oil cotton varieties.

Gossypium↗

Evidence for the prognostic value of TP53 mutations in circulating tumor DNA across solid malignancies: a systematic review and meta-analysis.

BACKGROUND: The purpose of this meta-analysis study is to provide evidence for the clinical utility of TP53 mutations in circulating tumor DNA (ctDNA) as a prognostic biomarker. METHODS: We searched the PubMed, Embase, Cochrane, and Web of Science databases (last update May 2025) for studies on TP53 mutations in ctDNA or cfDNA as prognosis overall survival and in solid tumors. A total of 21 studies that met the criteria were utilized and data was collected regarding the authors, year of publication, study design, site of the study, number of patients, detection, mutation sample size and outcome measures were collected. The Newcastle-Ottawa Scale (NOS) was used to evaluate the quality of the study, and meta-analysis was done by using STATA 16.0. Effect sizes were in the form of hazard ratios (HR) that had 95% confidence intervals (CI). The models used were fixed-effects and random-effects based on heterogeneity. Funnel plots, and Egger's test was used to measure publication bias, and sensitivity analysis conducted through a leave-one-out method. RESULTS: A total of 21 studies (2,685 TP53-mutated patients, one unreported) showed: Mutated patients had worse progression-free survival (PFS) (HR=2.10, p=0.000; 12 studies, heterogeneity resolved after excluding Yoshida 2023), shorter OS (HR=1.74, p=0.014; 9 studies), and reduced DFS (HR=1.73, p=0.007; 3 studies), but RFS (2 items) showed no statistically significant differences. Subgroup analyses revealed: Prospective studies showed stronger PFS (HR=2.14 vs retrospective 1.90) with Japanese subgroup HR=4.90; Lung/liver cancers had higher HRs than breast. Prospective OS HR=2.25 (lung 3.14, endometrial 0.75). Retrospective DFS HR=1.89 vs Japanese breast RFS HR=4.00. Heterogeneity originated from study design, region, and cancer type variations, with no significant publication bias (Egger's test p>0.05). CONCLUSION: Current evidence suggests that TP53 mutations detected in ctDNA are significantly associated with poor prognosis in various solid tumors, particularly lung cancer. The association is robust for PFS and OS, though high heterogeneity and biological complexity warrant cautious interpretation. These findings support the potential incorporation of ctDNA-based TP53 mutation status into clinical prognostic assessment systems as an adjunctive parameter; however, further standardization of detection protocols, functional annotation of mutation types (e.g., LOF vs. GOF), incorporation of VAF and clonality analysis, and validation in large prospective multicenter cohorts are needed before routine clinical implementation. PROSPERO REGISTRATION NUMBER: CRD420251021095.

Humans↗

Characterization of the gut phageome and functional genes carried by phages in laying hens with fatty liver hemorrhagic syndrome.

BACKGROUND: The gut microbiota is closely associated with the development of fatty liver hemorrhagic syndrome (FLHS); however, the function of its viral component, particularly bacteriophages, remains poorly understood. This study compared clinical parameters and the cecal phageome between 30-week-old (W30) and 50-week-old (W50) laying hens to characterize gut phages in the context of this metabolic disorder. RESULTS: Clinical analysis revealed that the W50 group exhibited typical FLHS, accompanied by elevated serum liver function and lipid markers (P&#x2009;<&#x2009;0.05). Functional prediction of the gut microbiota suggested a reduced lipid-metabolic capacity in W50 compared to the W30 group. A total of 20,274 phage genomes were identified from the two groups. These phages were primarily classified into 67 viral families, including Salasmaviridae, Herelleviridae, Suoliviridae, Peduoviridae, Crevaviridae, and Casjensviridae. The families Druskaviridae, Felixviridae, and Stanwilliamsviridae were uniquely detected in the W50 group. The phage community structure differed significantly between groups, with both phage diversity and richness markedly lower in W50 (P&#x2009;<&#x2009;0.05). LEfSe analysis revealed that phage taxa such as Stegnyidae, Herpelidae, and Chasovidae were significantly enriched in the W50 group, whereas Crewdviridae, Salasmaviridae, and Castroviridae were predominantly enriched in the W30 group. Functional annotation showed that these phages encode numerous metabolism-related genes and carry antimicrobial resistance genes (ARGs) as well as virulence factor genes. Notably, the diversity of ARGs carried by W50 phages was significantly higher (P&#x2009;<&#x2009;0.05), and ARG-rank analysis indicated a greater potential risk to human health. CONCLUSIONS: This study provides the first characterization of the gut phageome associated with FLHS in laying hens and confirms that gut phages constitute an important reservoir of ARGs. These findings offer a new perspective for understanding the pathogenesis of this disease and its associated public health risks. Video Abstract.

Animals↗

Application of bioinformatics in cancer epigenetics.

With the completion of the human genome sequence and the advent of high-throughput genomics-based technologies, it is now possible to study the entire human genome and epigenome. The challenge in the next decade of biomedical research is to functionally annotate the genome, epigenome, transcriptome, and proteome. High-throughput genome technology has already produced massive amounts of data including genome sequences, single nucleotide polymorphisms, and microarray gene expression. Our ability to manage and analyze data needs to match the speed of data acquisition. We will summarize our studies of allele-specific gene expression using genomic and computational approaches and identification of sequence motifs that are signature of imprinted genes. We will also discuss about how bioinformatics can facilitate epigenetic researches.

Computational Biology↗

Leiomyoma and myometrial gene expression profiles and their responses to gonadotropin-releasing hormone analog therapy.

Gene microarray was used to characterize the molecular environment of leiomyoma and matched myometrium during growth and in response to GnRH analog (GnRHa) therapy as well as GnRHa direct action on primary cultures of leiomyoma and myometrial smooth muscle cells (LSMC and MSMC). Unsupervised and supervised analysis of gene expression values and statistical analysis in R programming with a false discovery rate of P < or = 0.02 resulted in identification of 153 and 122 differentially expressed genes in leiomyoma and myometrium in untreated and GnRHa-treated cohorts, respectively. The expression of 170 and 164 genes was affected by GnRHa therapy in these tissues compared with their respective untreated group. GnRHa (0.1 microm), in a time-dependent manner (2, 6, and 12 h), targeted the expression of 281 genes (P < or = 0.005) in LSMC and MSMC, 48 of which genes were found in common with GnRHa-treated tissues. Functional annotations assigned these genes as key regulators of processes involving transcription, translational, signal transduction, structural activities, and apoptosis. We validated the expression of IL-11, early growth response 3, TGF-beta-induced factor, TGF-beta-inducible early gene response, CITED2 (cAMP response element binding protein-binding protein/p300-interacting transactivator with ED-rich tail), Nur77, growth arrest-specific 1, p27, p57, and G protein-coupled receptor kinase 5, representing cytokine, common transcription factors, cell cycle regulators, and signal transduction, at tissue levels and in LSMC and MSMC in response to GnRHa time-dependent action using real-time PCR, Western blotting, and immunohistochemistry. In conclusion, using different, complementary approaches, we characterized leiomyoma and myometrium molecular fingerprints and identified several previously unrecognized genes as targets of GnRHa action, implying that local expression and activation of these genes may represent features differentiating leiomyoma and myometrial environments during growth and GnRHa-induced regression.

Active Transport, Cell Nucleus↗

Isoflurane modulates genomic expression in rat amygdala.

General anesthesia, at a minimum, provides amnesia and unresponsiveness. Although anesthetics have many modulatory effects on neuronal ionophore protein complexes, it is not clear that the resulting electrophysiologic changes are the sole mechanisms of clinical anesthetic action. Cells respond to environmental changes in several ways, including alterations in DNA transcription leading to changes in the cell's proteins. We sought to expose the changes in global genomic expression, seeking potential targets involved in the processes of anesthetic-induced amnesia, and persistent long-term side effects of general anesthesia, including nausea and postoperative cognitive decline. Using Affymetrix GeneChips, we surveyed changes in expression across the entire expressed genome of Sprague-Dawley rat (n = 10 baseline, n = 6 isoflurane) basolateral amygdala 6 h after exposure to 15 min of 2% (1.4 MAC) isoflurane. Isoflurane administration was associated with altered expression in 269 unique genes possessing functional annotation. Affected genes were related to DNA transcription, protein synthesis, metabolism, signaling cascades, cytoskeletal structural proteins, and neural-specific proteins, among others. Even brief exposure to isoflurane leads to widespread changes in the genetic control in the amygdala 6 h after exposure. Gene expression is a dynamic process that may explain some long-term effects of anesthesia and that has the potential to modulate some of those effects using specific molecular therapeutics.

Amygdala↗

Intrinsic disorder is a common feature of hub proteins from four eukaryotic interactomes.

Recent proteome-wide screening approaches have provided a wealth of information about interacting proteins in various organisms. To test for a potential association between protein connectivity and the amount of predicted structural disorder, the disorder propensities of proteins with various numbers of interacting partners from four eukaryotic organisms (Caenorhabditis elegans, Saccharomyces cerevisiae, Drosophila melanogaster, and Homo sapiens) were investigated. The results of PONDR VL-XT disorder analysis show that for all four studied organisms, hub proteins, defined here as those that interact with > or = 10 partners, are significantly more disordered than end proteins, defined here as those that interact with just one partner. The proportion of predicted disordered residues, the average disorder score, and the number of predicted disordered regions of various lengths were higher overall in hubs than in ends. A binary classification of hubs and ends into ordered and disordered subclasses using the consensus prediction method showed a significant enrichment of wholly disordered proteins and a significant depletion of wholly ordered proteins in hubs relative to ends in worm, fly, and human. The functional annotation of yeast hubs and ends using GO categories and the correlation of these annotations with disorder predictions demonstrate that proteins with regulation, transcription, and development annotations are enriched in disorder, whereas proteins with catalytic activity, transport, and membrane localization annotations are depleted in disorder. The results of this study demonstrate that intrinsic structural disorder is a distinctive and common characteristic of eukaryotic hub proteins, and that disorder may serve as a determinant of protein interactivity.

Amino Acids↗