Search PubMed⌕ Search

Biomedical subjects

Falk Schubert

Publications and source records attributed to Falk Schubert.

9 recordsLinked to original sources

GOPET: a tool for automated predictions of Gene Ontology terms.

BACKGROUND: Vast progress in sequencing projects has called for annotation on a large scale. A Number of methods have been developed to address this challenging task. These methods, however, either apply to specific subsets, or their predictions are not formalised, or they do not provide precise confidence values for their predictions. DESCRIPTION: We recently established a learning system for automated annotation, trained with a broad variety of different organisms to predict the standardised annotation terms from Gene Ontology (GO). Now, this method has been made available to the public via our web-service GOPET (Gene Ontology term Prediction and Evaluation Tool). It supplies annotation for sequences of any organism. For each predicted term an appropriate confidence value is provided. The basic method had been developed for predicting molecular function GO-terms. It is now expanded to predict biological process terms. This web service is available via http://genius.embnet.dkfz-heidelberg.de/menu/biounit/open-husar CONCLUSION: Our web service gives experimental researchers as well as the bioinformatics community a valuable sequence annotation device. Additionally, GOPET also provides less significant annotation data which may serve as an extended discovery platform for the user.

Artificial Intelligence↗

High-resolution genomic profiling reveals association of chromosomal aberrations on 1q and 16p with histologic and genetic subgroups of invasive breast cancer.

PURPOSE: Invasive ductal carcinoma and invasive lobular carcinoma (ILC) represent the major histologic subtypes of invasive breast cancer. They differ with regard to presentation, metastatic spread, and epidemiologic features. To elucidate the genetic basis of these differences, we analyzed copy number imbalances that differentiate the histologic subtypes. EXPERIMENTAL DESIGN: High-resolution genomic profiling of 40 invasive breast cancers using matrix-comparative genomic hybridization with an average resolution of 0.5 Mb was conducted on bacterial artificial chromosome microarrays. The data were subjected to classification and unsupervised hierarchical cluster analyses. Expression of candidate genes was analyzed in tumor samples. RESULTS: The highest discriminating power was achieved when combining the aberration patterns of chromosome arms 1q and 16p, which were significantly more often gained in ILC. These regions were further narrowed down to subregions 1q24.2-25.1, 1q25.3-q31.3, and 16p11.2. Located within the candidate gains on 1q are two genes, FMO2 and PTGS2, known to be overexpressed in ILC relative to invasive ductal carcinoma. Assessment of four candidate genes on 16p11.2 by real-time quantitative PCR revealed significant overexpression of FUS and ITGAX in ILC with 16p copy number gain. Unsupervised hierarchical cluster analysis identified three molecular subgroups that are characterized by different aberration patterns, in particular concerning gain of MYC (8q24) and the identified candidate regions on 1q24.2-25.1, 1q25.3-q31.3, and 16p11.2. These genetic subgroups differed with regard to histology, tumor grading, frequency of alterations, and estrogen receptor expression. CONCLUSIONS: Molecular profiling using bacterial artificial chromosome arrays identified DNA copy number imbalances on 1q and 16p as significant classifiers of histologic and molecular subgroups.

Biomarkers, Tumor↗

Design optimization methods for genomic DNA tiling arrays.

A recent development in microarray research entails the unbiased coverage, or tiling, of genomic DNA for the large-scale identification of transcribed sequences and regulatory elements. A central issue in designing tiling arrays is that of arriving at a single-copy tile path, as significant sequence cross-hybridization can result from the presence of non-unique probes on the array. Due to the fragmentation of genomic DNA caused by the widespread distribution of repetitive elements, the problem of obtaining adequate sequence coverage increases with the sizes of subsequence tiles that are to be included in the design. This becomes increasingly problematic when considering complex eukaryotic genomes that contain many thousands of interspersed repeats. The general problem of sequence tiling can be framed as finding an optimal partitioning of non-repetitive subsequences over a prescribed range of tile sizes, on a DNA sequence comprising repetitive and non-repetitive regions. Exact solutions to the tiling problem become computationally infeasible when applied to large genomes, but successive optimizations are developed that allow their practical implementation. These include an efficient method for determining the degree of similarity of many oligonucleotide sequences over large genomes, and two algorithms for finding an optimal tile path composed of longer sequence tiles. The first algorithm, a dynamic programming approach, finds an optimal tiling in linear time and space; the second applies a heuristic search to reduce the space complexity to a constant requirement. A Web resource has also been developed, accessible at http://tiling.gersteinlab.org, to generate optimal tile paths from user-provided DNA sequences.

Algorithms↗

CGH-Profiler: data mining based on genomic aberration profiles.

BACKGROUND: CGH-Profiler is a program that supports the analysis of genomic aberrations measured by Comparative Genomic Hybridisation (CGH). Comparative genomic hybridisation (CGH) is a well-established, molecular cytogenetic method that allows the detection of chromosomal imbalances in entire genomes. This technique is widely used in routine molecular diagnostics. Typically, chromosomal imbalances are described in a complex syntax based on the International Standard for Cytogenetic Nomenclature (ISCN). This semantic description of chromosomal imbalances hinders a large-scale statistical analysis across different experiments, e.g. for finding aberration patterns associated with a particular disease type or state. RESULTS: CGH-Profiler circumvents the semantic ISCN description by importing data from different CGH system vendors and by directly transferring the data into a table format that is readily accessible for subsequent statistical analysis. CGH-profiler comes with different consistency checks, calculates various statistics and automatically assigns a median copy number ratio to each chromosomal band. Import of CGH profiles from different CGH system vendors is already supported; its extension to other systems can be readily achieved through Perl scripts.CGH profiler can also be used to analyse comparative expressed sequence hybridisation (CESH) data. CESH reveals gene expression patterns according to chromosomal locations in a similar manner as CGH detects chromosomal imbalances. CONCLUSION: CGH-Profiler is a useful tool for processing of CGH and CESH data.

Chromosome Aberrations↗

Bi-specific immunomagnetic enrichment of micrometastatic tumour cell clusters from bone marrow of cancer patients.

Metastasis-the spread of tumour cells from a primary lesion to distant organs-is the main cause of cancer-related death, and bone marrow (BM) is a frequent site for the settlement of disseminated tumour cells. Many BM samples harbour isolated tumour cells, whereas tumour cell clusters, as the potential precursors of solid distant metastases, are rarely detected after current enrichment procedures. We have analysed BM samples from 43 patients with carcinomas of the breast, colon and ovaries; 41 of these patients had no clinical signs of overt metastases (stage M0). Tumour cells in BM were enriched with immunomagnetic beads coupled to monoclonal antibodies against both EpCAM and HER2/neu. After enrichment, tumour cells were identified by immunostaining with the anti-cytokeratin antibody A45-B/B3. In total, 886 CK-positive cells were detected in 16 (35%) samples after immunomagnetic enrichment as compared to 34 cells in 9 (21%) samples using Ficoll density centrifugation previously used as the standard enrichment technique. Most remarkably, clusters of 2 to 10 CK-positive cells were found in 75% of CK-positive samples enriched by immunobeads, whereas no CK-positive cell clusters were detected after Ficoll enrichment. The method described offers an excellent tool for the enrichment of micrometastatic tumour cell clusters; these clusters may represent the initial stage of development from a single disseminated tumour cell towards an overt metastasis.

Bone Marrow Neoplasms↗

Genomic analysis of single cytokeratin-positive cells from bone marrow reveals early mutational events in breast cancer.

Chromosomal instability in human breast cancer is known to take place before mammary neoplasias display morphological signs of invasion. We describe here the unexpected finding of a tumor cell population with normal karyotypes isolated from bone marrow of breast cancer patients. By analyzing the same single cells for chromosomal aberrations, subchromosomal allelic losses, and gene amplifications, we confirmed their malignant origin and delineated the sequence of genomic events during breast cancer progression. On this trajectory of genomic progression, we identified a subpopulation of patients with very early HER2 amplification. Because early changes have the highest probability of being shared by genetically unstable tumor cells, the genetic characterization of disseminated tumor cells provides a novel rationale for selecting patients for targeted therapies.

Apoptosis↗

Applying Support Vector Machines for Gene Ontology based gene function prediction.

BACKGROUND: The current progress in sequencing projects calls for rapid, reliable and accurate function assignments of gene products. A variety of methods has been designed to annotate sequences on a large scale. However, these methods can either only be applied for specific subsets, or their results are not formalised, or they do not provide precise confidence estimates for their predictions. RESULTS: We have developed a large-scale annotation system that tackles all of these shortcomings. In our approach, annotation was provided through Gene Ontology terms by applying multiple Support Vector Machines (SVM) for the classification of correct and false predictions. The general performance of the system was benchmarked with a large dataset. An organism-wise cross-validation was performed to define confidence estimates, resulting in an average precision of 80% for 74% of all test sequences. The validation results show that the prediction performance was organism-independent and could reproduce the annotation of other automated systems as well as high-quality manual annotations. We applied our trained classification system to Xenopus laevis sequences, yielding functional annotation for more than half of the known expressed genome. Compared to the currently available annotation, we provided more than twice the number of contigs with good quality annotation, and additionally we assigned a confidence value to each predicted GO term. CONCLUSIONS: We present a complete automated annotation system that overcomes many of the usual problems by applying a controlled vocabulary of Gene Ontology and an established classification method on large and well-described sequence data sets. In a case study, the function for Xenopus laevis contig sequences was predicted and the results are publicly available at ftp://genome.dkfz-heidelberg.de/pub/agd/gene_association.agd_Xenopus.

Animals↗

Quality markers of drug information on the Internet: an evaluation of sites about St. John's wort.

PURPOSE: We aimed to evaluate websites about St. John's wort for the quality of their content, including accuracy as reflected by statement of correct indication and mentioning of interacting drugs, the presence of formal criteria as reflected by adherence to published standards for health information on the Internet, and the validity of individual formal criteria as markers of content quality. SUBJECTS AND METHODS: The Internet was searched with the metasearch engine WebFerret for sites about St. John's wort. A cross-sectional survey of a representative sample of randomly selected sites (n = 208) was performed. The main outcomes were the percentage of sites fulfilling the two criteria of content quality, the percentage of sites exhibiting eight formal criteria, and the associations between formal criteria and criteria of content quality, as determined by multivariate logistic regression analysis. RESULTS: Twenty-two percent (n = 45) of the websites correctly listed depression as the only indication for use of St. John's wort, and 22% (n = 46) identified at least one drug interaction with St. John's wort. Citing scientific publications was associated with mentioning the correct indication (odds ratio [OR] = 4.4; 95% confidence interval [CI]: 1.4 to 14) and mentioning of any interacting drug (OR = 6.0; 95% CI: 2.0 to 18). Absence of financial interest was associated with mentioning the correct indication (OR = 5.4; 95% CI: 2.2 to 14) and interacting drugs (OR = 3.1; 95% CI: 1.2 to 7.7). CONCLUSION: The content quality of sites about St. John's wort was generally poor. Our results suggest that Internet users should prefer noncommercial sites that reference the information to scientific publications when searching for drug information.

Cross-Sectional Studies↗

Microarray-based copy number and expression profiling in dedifferentiated and pleomorphic liposarcoma.

Sixteen dedifferentiated and pleomorphic liposarcomas were analyzed by comparative genomic hybridization (CGH) to genomic microarrays (matrix-CGH), cDNA-derived microarrays for expression profiling, and by quantitative PCR. Matrix-CGH revealed copy number gains of numerous oncogenes, i.e., CCND1, MDM2, GLI, CDK4, MYB, ESR1, and AIB1, several of which correlate with a high level of transcripts from the respective gene. In addition, a number of genes were found differentially expressed in dedifferentiated and pleomorphic liposarcomas. Application of dedicated clustering algorithms revealed that both tumor subtypes are clearly separated by the genomic profiles but only with a lesser power by the expression profiles. Using a support vector machine, a subset of five clones was identified as "class discriminators." Thus, for the distinction of these types of liposarcomas, genomic profiling appears to be more advantageous than RNA expression analysis.

Algorithms↗