Search PubMed⌕ Search

Biomedical subjects

Martin Vingron

Publications and source records attributed to Martin Vingron.

14 recordsLinked to original sources

Genome-wide array analysis of normal and malformed human hearts.

BACKGROUND: We present the first genome-wide cDNA array analysis of human congenitally malformed hearts and attempted to partially elucidate these complex phenotypes. Most congenital heart defects, which account for the largest number of birth defects in humans, represent complex genetic disorders. As a consequence of the malformation, abnormal hemodynamic features occur and cause an adaptation process of the heart. METHODS AND RESULTS: The statistical analysis of our data suggests distinct gene expression profiles associated with tetralogy of Fallot, ventricular septal defect, and right ventricular hypertrophy. Applying correspondence analysis, we could associate specific gene functions to specific phenotypes. Furthermore, our study design allows the suggestion that alterations associated with primary genetic abnormalities can be distinguished from those associated with the adaptive response of the heart to the malformation (right ventricular pressure overload hypertrophy). We provide evidence for the molecular transition of the hypertrophic right ventricle to normal left ventricular characteristics. Furthermore, we present data on chamber-specific gene expression. CONCLUSIONS: Our findings propose that array analysis of malformed human hearts opens a new window to understand the complex genetic network of cardiac development and adaptation. For detailed access, see the online-only Data Supplement.

Adaptation, Physiological↗

Gene expression profile of mouse bone marrow stromal cells determined by cDNA microarray analysis.

Bone marrow stromal cells (BMSC) have gained increased attention because of their multipotency and adult stem cell character. They have been shown to differentiate into other cell types of the mesenchymal lineage and also into non-mesenchymal cells. The exact identity of the original cells, which are isolated from bone marrow by their selective adherence to plastic, remains unknown to date. We have established and characterized mouse BMSC cultures and analyzed three independent samples by cDNA microarrays. The expression profile was compared with two previous expression studies of human BMSC and revealed a high degree of concordance between different techniques and species. To gain clues about the positional context and biology of the isolated cells within the bone marrow stroma, we searched our data for genes that encode proteins of the extracellular matrix, cell adhesion proteins, cytoskeletal proteins and cytokines/cytokine receptors. This analysis revealed a close association of BMSC with vascular cells and indicated that BMSC resemble pericytes.

Animals↗

Increase of functional diversity by alternative splicing.

A large-scale analysis of protein isoforms arising from alternative splicing shows that alternative splicing tends to insert or delete complete protein domains more frequently than expected by chance, whereas disruption of domains and other structural modules is less frequent. If domain regions are disrupted, the functional effect, as predicted from 3D structure, is frequently equivalent to removal of the entire domain. Also, short alternative splicing events within domains, which might preserve folded structure, target functional residues more frequently than expected. Thus, it seems that positive selection has had a major role in the evolution of alternative splicing.

Alternative Splicing↗

New evidence for genome-wide duplications at the origin of vertebrates using an amphioxus gene set and completed animal genomes.

The 2R hypothesis predicting two genome duplications at the origin of vertebrates is highly controversial. Studies published so far include limited sequence data from organisms close to the hypothesized genome duplications. Through the comparison of a gene catalog from amphioxus, the closest living invertebrate relative of vertebrates, to 3453 single-copy genes orthologous between Caenorhabditis elegans (C), Drosophila melanogaster (D), and Saccharomyces cerevisiae (Y), and to Ciona intestinalis ESTs, mouse, and human genes, we show with a large number of genes that the gene duplication activity is significantly higher after the separation of amphioxus and the vertebrate lineages, which we estimate at 650 million years (Myr). The majority of human orthologs of 195 CDY groups that could be dated by the molecular clock appear to be duplicated between 300 and 680 Myr with a mean at 488 million years ago (Mya). We detected 485 duplicated chromosomal segments in the human genome containing CDY orthologs, 331 of which are found duplicated in the mouse genome and within regions syntenic between human and mouse, indicating that these were generated earlier than the human-mouse split. Model based calculations of the codon substitution rate of the human genes included in these segments agree with the molecular clock duplication time-scale prediction. Our results favor at least one large duplication event at the origin of vertebrates, followed by smaller scale duplication closer to the bird-mammalian split.

Animals↗

SYSTERS, GeneNest, SpliceNest: exploring sequence space from genome to protein.

We have integrated the protein families from SYSTERS and the expressed sequence tag (EST) clusters from our database GeneNest with SpliceNest, a new database mapping EST contigs into genomic DNA. The SYSTERS protein sequence cluster set provides an automatically generated classification of all sequences of the SWISS-PROT, TrEMBL and PIR databases into disjoint protein family and superfamily clusters. GeneNest is a database and software package for producing and visualizing gene indices from ESTs and mRNAs. Currently, the database comprises gene indices of human, mouse, Arabidopsis thaliana and zebrafish. SpliceNest is a web-based graphical tool to explore gene structure, including alternative splicing, based on a mapping of the EST consensus sequences from GeneNest to the complete human genome. The integration of SYSTERS, GeneNest and SpliceNest into one framework now permits an overall exploration of the whole sequence space covering protein, mRNA and EST sequences, as well as genomic DNA. The databases are available for querying and browsing at http://cmb.molgen.mpg.de.

Alternative Splicing↗

Microarray data representation, annotation and storage.

Management and analysis of the huge amounts of data produced by microarray experiments is becoming one of the major bottlenecks in the utilization of this high-throughput technology. We describe the basic design of a microarray gene expression database to help microarray users and their informatics teams to set up their information services. We describe two data models--a simpler one called ArrayExpressB and the complete model ArrayExpressC, and discuss some implementation issues. For latest developments see http: wwwebi.ac.uk/arrayexpress

Algorithms↗

Microarray data warehouse allowing for inclusion of experiment annotations in statistical analysis.

MOTIVATION: Microarray technology provides access to expression levels of thousands of genes at once, producing large amounts of data. These datasets are valuable only if they are annotated by sufficiently detailed experiment descriptions. However, in many databases a substantial number of these annotations is in free-text format and not readily accessible to computer-aided analysis. RESULTS: The Multi-Conditional Hybridization Intensity Processing System (M-CHIPS), a data warehousing concept, focuses on providing both structure and algorithms suitable for statistical analysis of a microarray database's entire contents including the experiment annotations. It addresses the rapid growth of the amount of hybridization data, more detailed experimental descriptions, and new kinds of experiments in the future. We have developed a storage concept, a particular instance of which is an organism-specific database. Although these databases may contain different ontologies of experiment annotations, they share the same structure and therefore can be accessed by the very same statistical algorithms. Experiment ontologies have not yet reached their final shape, and standards are reduced to minimal conventions that do not yet warrant extensive description. An ontology-independent structure enables updates of annotation hierarchies during normal database operation without altering the structure. AVAILABILITY AND SUPPLEMENTARY INFORMATION: http://www.dkfz.de/tbi/services/mchips

Algorithms↗

TREE-PUZZLE: maximum likelihood phylogenetic analysis using quartets and parallel computing.

SUMMARY: TREE-PUZZLE is a program package for quartet-based maximum-likelihood phylogenetic analysis (formerly PUZZLE, Strimmer and von Haeseler, Mol. Biol. Evol., 13, 964-969, 1996) that provides methods for reconstruction, comparison, and testing of trees and models on DNA as well as protein sequences. To reduce waiting time for larger datasets the tree reconstruction part of the software has been parallelized using message passing that runs on clusters of workstations as well as parallel computers. AVAILABILITY: http://www.tree-puzzle.de. The program is written in ANSI C. TREE-PUZZLE can be run on UNIX, Windows and Mac systems, including Mac OS X. To run the parallel version of PUZZLE, a Message Passing Interface (MPI) library has to be installed on the system. Free MPI implementations are available on the Web (cf. http://www.lam-mpi.org/mpi/implementations/).

Algorithms↗

Variance stabilization applied to microarray data calibration and to the quantification of differential expression.

We introduce a statistical model for microarray gene expression data that comprises data calibration, the quantification of differential expression, and the quantification of measurement error. In particular, we derive a transformation h for intensity measurements, and a difference statistic Deltah whose variance is approximately constant along the whole intensity range. This forms a basis for statistical inference from microarray data, and provides a rational data pre-processing strategy for multivariate analyses. For the transformation h, the parametric form h(x)=arsinh(a+bx) is derived from a model of the variance-versus-mean dependence for microarray intensity data, using the method of variance stabilizing transformations. For large intensities, h coincides with the logarithmic transformation, and Deltah with the log-ratio. The parameters of h together with those of the calibration between experiments are estimated with a robust variant of maximum-likelihood estimation. We demonstrate our approach on data sets from different experimental platforms, including two-colour cDNA arrays and a series of Affymetrix oligonucleotide arrays.

Algorithms↗

Theoretical analysis of alternative splice forms using computational methods.

Nowadays understanding alternative splicing is one of the greatest challenges in biology, because it is a genetic process much more important than thought at the time of its discovery. In this paper, we explain the approach of using the different available databases and software tools to start a large scale investigation of alternative splice forms. To collect information about alternative splicing we investigated known data in the databases using different computational methods. The investigations proceeded from the genomic sequence data to structural protein data. Then, we interpreted those data to find the relationship between alternative splice forms and protein function and structure. We found some interesting features of alternative splicing which are presented here. We discuss the results of one chosen example. They concern the coverage quality of the protein sequence of a known structure, an EST analysis, the validation of splice variants, the determination of the alternative splice type, and finally the link between alternative splicing and disease.

Algorithms↗

Annotating regulatory DNA based on man-mouse genomic comparison.

Non-coding DNA segments that are conserved between the human and mouse genomic sequence are good indicators of possible regulatory sequences. Here we report on a systematic approach to delineate such conserved elements from upstream regions of orthologous gene pairs from man and mouse. We focus on orthologous genes in order to maximize our chances to find functionally similar regulatory elements. The identification of conserved elements is effected using the Waterman-Eggert local suboptimal alignment algorithm. We have modified an implementation of this algorithm such that it integrates the determination of statistical significance for the local suboptimal alignments. This has the effect of outputting a dynamically determined number of suboptimal alignments that are deemed statistically significant. Comparison with experimentally determined annotation shows a striking enrichement of regulatory sites among the conserved regions. Furthermore, the conserved regions tend to cover the promotor region described in the EPD database.

Algorithms↗

Estimating amino acid substitution models: a comparison of Dayhoff's estimator, the resolvent approach and a maximum likelihood method.

Evolution of proteins is generally modeled as a Markov process acting on each site of the sequence. Replacement frequencies need to be estimated based on sequence alignments. Here we compare three approaches: First, the original method by Dayhoff, Schwartz, and Orcutt (1978) Atlas Protein Seq. Struc. 5:345-352, secondly, the resolvent method (RV) by Müller and Vingron (2000) J. Comput. Biol. 7(6):761-776, and finally a maximum likelihood approach (ML) developed in this paper. We evaluate the methods using a highly divergent and inhomogeneous set of sequence alignments as an input to the estimation procedure. ML is the method of choice for small sets of input data. Although the RV method is computationally much less demanding it performs only slightly worse than ML. Therefore, it is perfectly appropriate for large-scale applications.

Algorithms↗

Monitoring the switch from housekeeping to pathogen defense metabolism in Arabidopsis thaliana using cDNA arrays.

Plants respond to pathogen attack by deploying several defense reactions. Some rely on the activation of preformed components, whereas others depend on changes in transcriptional activity. Using cDNA arrays comprising 13,000 unique expressed sequence tags, changes in the transcriptome of Arabidopsis thaliana were monitored after attempted infection with the bacterial plant pathogen Pseudomonas syringae pv. tomato carrying the avirulence gene avrRpt2. Sampling at four time points during the first 24 h after infiltration revealed significant changes in the steady state transcript levels of approximately 650 genes within 10 min and a massive shift in gene expression patterns by 7 h involving approximately 2,000 genes representing many cellular processes. This shift from housekeeping to defense metabolism results from changes in regulatory and signaling circuits and from an increased demand for energy and biosynthetic capacity in plants fighting off a pathogenic attack. Concentrating our detailed analysis on the genes encoding enzymes in glycolysis, the Krebs cycle, the pentose phosphate pathway, the biosynthesis of aromatic amino acids, phenylpropanoids, and ethylene, we observed interesting differential regulation patterns. Furthermore, our data showed potentially important changes in areas of metabolism, such as the glyoxylate metabolism, hitherto not suspected to be components of plant defense.

Arabidopsis↗