Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Comprehensive genomic profiling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40Linked to original sources

Genetic diversity among Mycoplasma species bovine group 7: clonal isolates from an outbreak of polyarthritis, mastitis, and abortion in dairy cattle.

A comprehensive genetic analysis of 60 Mycoplasma sp. bovine group 7 isolates from different geographic origins and epidemiological settings is presented. Twenty-four isolates were recovered from the joints of calves during sporadic episodes of polyarthritis in geographically distinct regions of Queensland and New South Wales, Australia, including two clones of the type strain PG5O. A further three Australian isolates were also recovered from the tympanic bulla, retropharyngeal lymph node and the lung and another three isolates had unconfirmed histories. Six isolates originated from Germany, Portugal, Nigeria, and France. Twenty-four epidemiologically related isolates of Mycoplasma sp. bovine group 7 were recovered from multiple tissue sites and body fluids of infected calves with polyarthritis, mastitic milk, and from the stomach contents, lung and liver from aborted foetuses in three large, centrally managed dairy herds in New South Wales, Australia. Restriction endonuclease analysis (REA) of genomic DNA differentiated 29 Cfol profiles among these 60 isolates and grouped all 24 epidemiologically related isolates in a defined pattern showing a clonal origin. Three isolates of this clonal cluster were recovered from mastitic milk and the synovial exudate of clinically-affected calves and appeared sporadically for periods up to 18 months after the initial outbreak of polyarthritis indicating a persistent, close association of the organism with cattle in these herds. The Cfol profile representative of the clonal cluster was distinguishable from profiles of isolates recovered from multiple, unrelated cases of polyarthritis in Queensland and New South Wales and from other countries. All 24 isolates from the clonal cluster possessed a plasmid (pBG7AU) with a molecular size of 1022 bp. DNA sequence analysis of pBG7AU identified two open reading frames sharing 81 and 99% DNA sequence similarity with hypothetical replication control proteins A and B respectively, previously described in plasmid pADB201 isolated from M. mycoides subspecies mycoides. Other isolates of bovine group 7, epidemiologically unrelated to the clonal cluster, including two clones of the type strain PG5O, possessed a similar-sized plasmid. These data confirm that Mycoplasma sp. bovine group 7 is capable of migrating to, and multiplying within, different tissue sites within a single animal and among different animals within a herd.

Abortion, Veterinary↗

Proposal of real-world solutions for the implementation of predictive biomarker testing in patients with operable non-small cell lung cancer.

The implementation of biomarker testing for targeted therapies and immune checkpoint inhibitors is a cornerstone in the management of metastatic and locally advanced non-small cell lung cancer (NSCLC), playing a pivotal role in guiding treatment decisions and patient care. The emergence of precision medicine in the realm of operable NSCLC has been marked by the recent approvals of osimertinib, atezolizumab, nivolumab, pembrolizumab and alectinib for early-stage disease, signifying a shift towards more tailored therapeutic strategies. Concurrently, the landscape of this disease is rapidly evolving, with several further pending approvals and numerous clinical trials in progress. To harness the benefits of these innovative neo-adjuvant and adjuvant therapies, the integration of predictive biomarker testing into standard clinical protocols is imperative for patients with operable NSCLC. A multidisciplinary international consortium has identified three primary obstacles impeding the effective testing of patients with operable NSCLC. These challenges encompass the limited number of test requests by physicians, the inadequacy of tissue samples for comprehensive testing, and the prevalence of cost-reduction measures leading to suboptimal testing practices. This review delineates the aforementioned challenges and proposed solutions, and strategic recommendations aimed at enhancing the testing process. By addressing these issues, we strive to optimize patient outcomes in operable NSCLC, ensuring that individuals receive the most appropriate and effective care based on their unique disease profile.

Humans↗

Functional genomics of the social amoebae, Dictyostelium discoideum.

Dictyostelium discoideum is one of the simplest organisms to form a multicellular structure, and it offers several advantages as a model. In order to understand the genetic basis of the multicellular development, a comprehensive analysis of cDNAs is being performed. To date, about 75,000 ESTs have been collected at different stages of development. They have been assembled into about 6,400 independent sequences that represent 70-80% of all of the expected genes in D. discoideum. The results are available on the Internet. In addition to structural analyses, functional analyses of the temporal and spatial expression patterns and gene targeting are being carried out. Furthermore, there are plans to combine the information that is obtained from the cDNA, Genome, and Proteome Projects, as well as the published results, into an integrated database, DictyBase.

Animals↗

Identification of conserved pathways of DNA-damage response and radiation protection by genome-wide RNAi.

Ionizing radiation is extremely harmful for human cells, and DNA double-strand breaks (DSBs) are considered to be the main cytotoxic lesions induced. Improper processing of DSBs contributes to tumorigenesis, and mutations in DSB response genes underlie several inherited disorders characterized by cancer predisposition. Here, we performed a comprehensive screen for genes that protect animal cells against ionizing radiation. A total of 45 C. elegans genes were identified in a genome-wide RNA interference screen for increased sensitivity to ionizing radiation in germ cells. These genes include orthologs of well-known human cancer predisposition genes as well as novel genes, including human disease genes not previously linked to defective DNA-damage responses. Knockdown of eleven genes also impaired radiation-induced cell-cycle arrest, and seven genes were essential for apoptosis upon exposure to irradiation. The gene set was further clustered on the basis of increased sensitivity to DNA-damaging cancer drugs cisplatin and camptothecin. Almost all genes are conserved across animal phylogeny, and their relevance for humans was directly demonstrated by showing that their knockdown in human cells results in radiation sensitivity, indicating that this set of genes is important for future cancer profiling and drug development.

Animals↗

Molecular insights into the initiation of sporulation in Gram-positive bacteria: new technologies for an old phenomenon.

The last decade has witnessed extensive, and widespread, changes in scientific technologies that have impacted significantly upon the study of the life sciences. Arguably, the biggest advances in our comprehension of simple and complex biological processes have come as a consequence of obtaining the complete DNA sequence of organisms. It is likely that we will become accustomed to hearing of quantum leaps in the study and understanding of the biology of higher eukaryotes in the coming years, now that (near) complete genome sequences are available for man, mouse and rat. In this review, we will discuss the impact of genome sequence data, and the use of new scientific technologies that have emerged largely as consequence of the availability of this information, on the study of the master regulator of sporulation, Spo0A, in low G+C Gram-positive endospore-forming bacteria.

Bacterial Proteins↗

Overview and perspectives the transcriptome of Paracoccidioides brasiliensis.

Paracoccidioides brasiliensis is a dimorphic and thermo-regulated fungus which is the causative agent of paracoccidioidomycosis, an endemic disease widespread in Latin America that affects 10 million individuals. Pathogenicity is assumed to be a consequence of the dimorphic transition from mycelium to yeast cells during human infection. This review shows the results of the P. brasiliensis transcriptome project which generated 6,022 assembled groups from mycelium and yeast phases. Computer analysis using the tools of bioinformatics revealed several aspects from the transcriptome of this pathogen such as: general and differential metabolism in mycelium and yeast cells; cell cycle, DNA replication, repair and recombination; RNA biogenesis apparatus; translation and protein fate machineries; cell wall; hydrolytic enzymes; proteases; GPI-anchored proteins; molecular chaperones; insights into drug resistance and transporters; oxidative stress response and virulence. The present analysis has provided a more comprehensive view of some specific features considered relevant for the understanding of basic and applied knowledge of P. brasiliensis.

Cell Wall↗

Comprehensive evaluation of AlphaFold/OpenFold prediction of experimentally unresolved proteins through novel metrics.

Predicting accurate protein structures is essential for understanding molecular mechanisms, interpreting the impact of sequence variation, and supporting translational applications ranging from drug discovery to clinical genomics. Recent advances in deep-learning-based predictors such as AlphaFold2, OpenFold, and AlphaFold3 have transformed structural biology, enabling routine in silico modeling even for challenging or previously uncharacterized proteins. However, systematic benchmarking of these tools-especially for novel targets and single amino acid variants-remains limited. Conventional global metrics often fail to capture biologically meaningful discrepancies. By evaluating multiple implementations of AlphaFold2 and OpenFold, together with ColabFold and the AlphaFold3 server, across 10 different proteins and 222 single amino acid protein variants encompassing a wide range of sizes, structures, and functions, we show that although widely used global indicators-like mean pLDDT, pTM-score, and RMSD-frequently suggest comparable performance, substantial local-level differences remain elusive. To address this gap, we introduce a comparative framework leveraging Bland-Altman agreement analysis, to evaluate per-residue Cα-confidence differences and Per-Residue profiles (PRPs), complemented by Uniform Manifold Approximation and Projection (UMAP). This approach reveals marked localized divergences, particularly within flexible or intrinsically disordered regions, where both predictor choice and single-residue substitutions trigger the largest conformational shifts. We further demonstrate that using reduced homology databases has minimal impact on predicted structural quality, offering computationally efficient alternatives. Collectively, our findings underscore the importance of integrating global and residue-specific evaluations to more accurately assess robustness, agreement, and practical usability across contemporary protein structure prediction methods.

Proteins↗

Comprehensive assessment of homologous recombination deficiency via simultaneous methylation and mutation analysis in epithelial ovarian cancer: implications for PARP inhibitors efficacy.

BACKGROUND: The advent of poly (ADP-ribose) polymerase inhibitors (PARPi) over the past decade has significantly altered the management of epithelial ovarian cancer (EOC). We proposed that the etiology of homologous recombination deficiency (HRD) might underlie the variable responses to PARPi observed across patient populations. METHODS: As part of the phase 2 study of the Chinese HRD Harmonization Project, we developed a genomic methylation sequencing (GM-seq) pipeline facilitated by the TET enzyme for the simultaneous identification of methylated modifications and genetic variations in EOC tumor samples, and compared with established DNA sequencing-based HRD assays. RESULTS: Somatic mutation and HRD scores were confounded by low tumor purity in our cohort of 98 locally advanced/advanced EOC patients. In samples with tumor purity&#x2009;&#x2265;&#x2009;30% (n&#x2009;=&#x2009;45), the GM-seq pipeline showed high consistency with DNA sequencing-based HRD assay, identifying genetic variations in homologous recombination repair (HRR) genes and HRD score with 92.6% (25/27) and 97.1% (33/34) consistency respectively, in addition to conducting methylation profiling. Moreover, different underlying mechanisms of HRD were associated with varying degrees of PARPi efficacy, with BRCA1/2 LOH group having the best efficacy (median PFS, undefined), followed by BRCA1 methylation group (median PFS, 23.4 months), and those with unknown etiology of HRD having the worst efficacy (median PFS, 8.8 months, p&#x2009;<&#x2009;0.001). CONCLUSION: Our findings underscore the importance of considering HRD etiology when evaluating PARPi efficacy in EOC patients. The GM-seq pipeline, represents a significant advancement in HRD detection, enabling more accurate predictions of PARPi response.

Epithelial ovarian cancer (EOC)↗

Genetic insights into lung squamous cell carcinoma: how TP53 and CSMD3 co-mutations shape prognosis and immune response.

BACKGROUND: Lung squamous cell carcinoma (LUSC) accounts for a significant proportion of lung cancer cases and is often associated with smoking and various environmental factors. The prognostic and immunologic implications of TP53 and CSMD3 co-mutations in LUSC remain poorly understood. This study aimed to investigate the role of TP53/CSMD3 co-mutations in LUSC using comprehensive bioinformatics analyses. METHODS: Data from 487 LUSC patients were obtained from The Cancer Genome Atlas (TCGA) database, with external validation performed using the combined cohort. Patients were stratified into TP53/CSMD3 co-mutation, single-mutation, and wild-type (WT) groups. Prognostic analysis was conducted using Kaplan-Meier survival curves. Tumor mutational burden (TMB) was calculated, and immune cell infiltration was assessed using multiple algorithms. Differentially expressed genes (DEGs) between co-mutated and WT groups were identified, followed by Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses. A nomogram incorporating mutation status, gender, age, and tumor stage (T stage) was developed for individualized prognostic prediction. RESULTS: The TP53/CSMD3 co-mutated group exhibited significantly better overall survival (OS) compared to single-mutation and WT groups. TMB scores were markedly higher in co-mutated patients, suggesting potential sensitivity to immune checkpoint inhibitors. Immune infiltration analysis revealed distinct profiles, including elevated CD8 T cells and reduced immunosuppressive components, in the co-mutation group. A total of 403 DEGs were identified between co-mutated and WT groups, with significant enrichment in immune-related pathways. Mechanistically, the co-mutation was associated with distinct downregulation of complement negative regulators (CFH/CFI), indicating complement hyperactivation independent of TMB. The constructed nomogram provided accurate individualized prognostic assessments. CONCLUSIONS: The co-mutation of TP53 and CSMD3 identifies a distinct LUSC subtype with favorable survival, marked by high TMB and an immune-activated microenvironment. Beyond TMB-driven neoantigen generation, the significant downregulation of complement negative regulators (CFH/CFI) reveals an independent complement hyperactivation pathway associated with CSMD3 loss. The constructed nomogram provides accurate individualized survival prediction. These findings establish TP53/CSMD3 co-mutation as a promising prognostic biomarker and offer mechanistic insights for personalized immunotherapy strategies. Future prospective cohorts are warranted to validate its predictive value.

Lung squamous cell carcinoma (LUSC)↗

Prognostic impact of CEACAM5 in nonsquamous non-small cell lung cancer: a critical reappraisal driven by cut-off optimization.

BACKGROUND: Carcinoembryonic antigen-related cell adhesion molecule 5 (CEACAM5) has emerged as a promising therapeutic target for antibody-drug conjugates (ADCs) in nonsquamous non-small cell lung cancer (NSCLC). However, the landscape of CEACAM5 protein expression and its association with clinicopathological features in nonsquamous NSCLC remain poorly characterized. This study represents the first comprehensive investigation to profile CEACAM5 protein expression in this specific population using a clinical-grade immunohistochemistry (IHC) assay. METHODS: We retrospectively analyzed 218 patients with resected nonsquamous NSCLC. A standardized IHC assay derived from a clinical trial was employed, utilizing a proprietary monoclonal antibody (clone 769) specifically developed for clinical application. We systematically evaluated the relationship between CEACAM5 expression and tumor characteristics using both a conventional therapeutic threshold (&#x2265;50%) and a sensitive data-driven cut-off (>0%) to explore the full spectrum of antigen expression. RESULTS: The assay demonstrated excellent reproducibility. We found that CEACAM5 expression was significantly associated with markers of tumor aggressiveness, including lymphovascular and pleural invasion. Notably, the sensitive >0% cut-off revealed broader biological correlations than the restrictive &#x2265;50% threshold. While CEACAM5 was not an independent prognostic factor in the overall cohort, a stage-specific, hypothesis-generating analysis-supported by external genomic validation-suggested that positive CEACAM5 expression may be associated with a trend towards poorer disease-free survival specifically in early-stage (stage I-II) patients, although this did not reach statistical significance (P=0.06). CONCLUSIONS: This study provides the first detailed characterization of CEACAM5 protein expression in nonsquamous NSCLC, establishing it as a biomarker linked to aggressive disease phenotypes. Our findings suggest that adopting a sensitive detection threshold (>0%) may better capture the population of patients with biologically active CEACAM5, thereby potentially refining candidate selection for targeted therapies and risk stratification in early-stage disease, a hypothesis that warrants prospective validation.

Carcinoembryonic antigen-related cell adhesion mol↗

Bacillus subtilis functional genomics: global characterization of the stringent response by proteome and transcriptome analysis.

The stringent response in Bacillus subtilis was characterized by using proteome and transcriptome approaches. Comparison of protein synthesis patterns of wild-type and relA mutant cells cultivated under conditions which provoke the stringent response revealed significant differences. According to their altered synthesis patterns in response to DL-norvaline, proteins were assigned to four distinct classes: (i) negative stringent control, i.e., strongly decreased protein synthesis in the wild type but not in the relA mutant (e.g., r-proteins); (ii) positive stringent control, i.e., induction of protein synthesis in the wild type only (e.g., YvyD and LeuD); (iii) proteins that were induced independently of RelA (e.g., YjcI); and (iv) proteins downregulated independently of RelA (e.g., glycolytic enzymes). Transcriptome studies based on DNA macroarray techniques were used to complement the proteome data, resulting in comparable induction and repression patterns of almost all corresponding genes. However, a comparison of both approaches revealed that only a subset of RelA-dependent genes or proteins was detectable by proteomics, demonstrating that the transcriptome approach allows a more comprehensive global gene expression profile analysis. The present study presents the first comprehensive description of the stringent response of a bacterial species and an almost complete map of protein-encoding genes affected by (p)ppGpp. The negative stringent control concerns reactions typical of growth and reproduction (ribosome synthesis, DNA synthesis, cell wall synthesis, etc.). Negatively controlled unknown y-genes may also code for proteins with a specific function during growth and reproduction (e.g., YlaG). On the other hand, many genes are induced in a RelA-dependent manner, including genes coding for already-known and as-yet-unknown proteins. A passive model is preferred to explain this positive control relying on the redistribution of the RNA polymerase under the influence of (p)ppGpp.

Bacillus subtilis↗

Highly multiplexed molecular inversion probe genotyping: over 10,000 targeted SNPs genotyped in a single tube assay.

Large-scale genetic studies are highly dependent on efficient and scalable multiplex SNP assays. In this study, we report the development of Molecular Inversion Probe technology with four-color, single array detection, applied to large-scale genotyping of up to 12,000 SNPs per reaction. While generating 38,429 SNP assays using this technology in a population of 30 trios from the Centre d'Etude Polymorphisme Humain family panel as part of the International HapMap project, we established SNP conversion rates of approximately 90% with concordance rates >99.6% and completeness levels >98% for assays multiplexed up to 12,000plex levels. Furthermore, these individual metrics can be "traded off" and, by sacrificing a small fraction of the conversion rate, the accuracy can be increased to very high levels. No loss of performance is seen when scaling from 6,000plex to 12,000plex assays, strongly validating the ability of the technology to suppress cross-reactivity at high multiplex levels. The results of this study demonstrate the suitability of this technology for comprehensive association studies that use targeted SNPs in indirect linkage disequilibrium studies or that directly screen for causative mutations.

Chromosome Inversion↗

Improved proteome coverage by using high efficiency cysteinyl peptide enrichment: the human mammary epithelial cell proteome.

Automated multidimensional capillary liquid chromatography-tandem mass spectrometry (LC-MS/MS) has been increasingly applied in various large scale proteome profiling efforts. However, comprehensive global proteome analysis remains technically challenging due to issues associated with sample complexity and dynamic range of protein abundances, which is particularly apparent in mammalian biological systems. We report here the application of a high efficiency cysteinyl peptide enrichment (CPE) approach to the global proteome analysis of human mammary epithelial cells (HMECs) which significantly improved both sequence coverage of protein identifications and the overall proteome coverage. The cysteinyl peptides were specifically enriched by using a thiol-specific covalent resin, fractionated by strong cation exchange chromatography, and subsequently analyzed by reversed-phase capillary LC-MS/MS. An HMEC tryptic digest without CPE was also fractionated and analyzed under the same conditions for comparison. The combined analyses of HMEC tryptic digests with and without CPE resulted in a total of 14 416 confidently identified peptides covering 4294 different proteins with an estimated 10% gene coverage of the human genome. By using the high efficiency CPE, an additional 1096 relatively low abundance proteins were identified, resulting in 34.3% increase in proteome coverage; 1390 proteins were observed with increased sequence coverage. Comparative protein distribution analyses revealed that the CPE method is not biased with regard to protein M(r) , pI, cellular location, or biological functions. These results demonstrate that the use of the CPE approach provides improved efficiency in comprehensive proteome-wide analyses of highly complex mammalian biological systems.

Amino Acid Sequence↗

Discovering protein-protein interactions.

The ongoing genomics and proteomics efforts have helped identify many new genes and proteins in living organisms. However, simply knowing the existence of genes and proteins does not tell us much about the biological processes in which they participate. Many major biological processes are controlled by protein interaction networks. A comprehensive description of protein-protein interactions is therefore necessary to understand the genetic program of life. In this tutorial, we provide an overview of the various current high-throughput methods for discovering protein-protein interactions, covering both the conventional experimental methods and new computational approaches.

Artificial Gene Fusion↗

Prospects offered by genome studies for combating meningococcal disease by vaccination.

Meningococcal disease was first recognised and Neisseria meningitidis isolated as the causative agent over 100 years ago, but despite more than a century of research, attempts to eliminate this distressing illness have so far been thwarted. The main problem lies in the fact that N. meningitidis usually exists as a harmless commensal inhabitant of the human nasopharynx, the pathogenic state being the exception rather than the norm. As man is its only host, the meningococcus is uniquely adapted to this ecological niche and has evolved an array of mechanisms for evading clearance by the human immune response. Progress has been made in combating the disease by developing vaccines that target specific pathogenic serogroups of meningococci. However, a fully comprehensive vaccine that protects against all pathogenic strains is still just beyond reach. The publication of the genome sequences of two meningococcal strains, one each from serogroups A and B and the imminent completion of a third illustrates the extent of the problems to be overcome, namely the vast array of genetic mechanisms for the generation of meningococcal diversity. Fortunately, genome studies also provide new hope for solutions to these problems in the potential for a greater understanding of meningococcal pathogenesis and possibilities for the identification of new vaccine candidates. This review describes some of the approaches that are currently being used to exploit the information from meningococcal genome sequences and seeks to identify future prospects for combating meningococcal disease.

Clinical Trials as Topic↗

The MYB transcription factor superfamily of Arabidopsis: expression analysis and phylogenetic comparison with the rice MYB family.

MYB proteins are a superfamily of transcription factors that play regulatory roles in developmental processes and defense responses in plants. We identified 198 genes in the MYB superfamily from an analysis of the complete Arabidopsis genome sequence, among them, 126 are R2R3-MYB, 5 are R1R2R3-MYB, 64 are MYB-related, and 3 atypical MYB genes. Here we report the expression profiles of 163 genes in the Arabidopsis MYB superfamily whose full-length open reading frames have been isolated. This analysis indicated that the expression for most of the Arabidopsis MYB genes were responsive to one or multiple types of hormone and stress treatments. A phylogenetic comparison of the members of this superfamily in Arabidopsis and rice suggested that the Arabidopsis MYB superfamily underwent a rapid expansion after its divergence from monocots but before its divergence from other dicots. It is likely that the MYB-related family was more ancient than the R2R3-MYB gene family, or had evolved more rapidly. Therefore, the MYB gene superfamily represents an excellent system for investigating the evolution of large and complex gene families in higher plants. Our comprehensive analysis of this largest transcription factor superfamily of Arabidopsis and rice may help elucidate the possible biological roles of the MYB genes in various aspects of flowering plants.

Amino Acid Sequence↗

Comprehensive identification and characterization of diallelic insertion-deletion polymorphisms in 330 human candidate genes.

Despite being the second most frequent type of polymorphism in the genome, diallelic insertion-deletion polymorphisms (indels) have received far less attention in the study of sequence variation. In this report, we describe an approach that can detect indels in the heterozygous state and can comprehensively identify indels in the target sequence. Using this approach, we identified 2393 indels in a set of 330 candidate genes, i.e. an average of seven indels per gene with about two indels per gene being common (minor allele frequency >or=0.1). We compared the population genetic characteristics of indels with substitutions in this data. Our data supported the findings that deletions occur more frequently in the human genome. 5'-UTR and coding regions of the genes showed a significantly lower diversity for indels compared with other regions, suggesting differences in effects of selection on indels and substitutions. Sequence diversity and pairwise linkage disequilibrium (LD) findings of the different populations were similar to earlier results and included a greater skew towards low-frequency variants and a faster rate of LD decay in the African-descent population compared with the non-African populations. Within populations, the allele frequency spectra and LD-decay profiles for indels were similar to substitutions. Overall, the findings suggest that, although the mechanisms giving rise to indels may be different from those causing substitutions, the evolutionary histories of indels and substitutions are similar, and that indels can play a valuable role in association studies and marker selection strategies.

5' Untranslated Regions↗

Improving gene annotation using peptide mass spectrometry.

Annotation of protein-coding genes is a key goal of genome sequencing projects. In spite of tremendous recent advances in computational gene finding, comprehensive annotation remains a challenge. Peptide mass spectrometry is a powerful tool for researching the dynamic proteome and suggests an attractive approach to discover and validate protein-coding genes. We present algorithms to construct and efficiently search spectra against a genomic database, with no prior knowledge of encoded proteins. By searching a corpus of 18.5 million tandem mass spectra (MS/MS) from human proteomic samples, we validate 39,000 exons and 11,000 introns at the level of translation. We present translation-level evidence for novel or extended exons in 16 genes, confirm translation of 224 hypothetical proteins, and discover or confirm over 40 alternative splicing events. Polymorphisms are efficiently encoded in our database, allowing us to observe variant alleles for 308 coding SNPs. Finally, we demonstrate the use of mass spectrometry to improve automated gene prediction, adding 800 correct exons to our predictions using a simple rescoring strategy. Our results demonstrate that proteomic profiling should play a role in any genome sequencing project.

Algorithms↗