Molecular genetics and genetic epidemiology of cardiovascular disease and diabetes. Introductory remarks: genetic models and statistical approaches.
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
A data base of gametic distributions at a stable equilibrium for genetic systems with up to five diallelic loci was created by numerically iterating equations for the dynamics of gametic frequencies in multilocus systems under selection. For a given number of loci, iterations were conducted for 4000 random sets of genotypic fitnesses, 6 values of recombination, and 10 different initial distributions. The data base was used to investigate the following properties of stable equilibria maintaining a polymorphism in a given number of loci that are expected a priori, i.e., without any constraints on fitnesses of genotypes: probability for a fitness set to yield a such equilibrium; probability for a random trajectory to converge to a such equilibrium; genetic load at a such equilibrium. The expected number of simultaneously stable equilibria, and the fraction of genome maintained polymorphic were also investigated as well as some parameters expected at an equilibrium maintaining all loci polymorphic. One of the most important findings is that multilocus genetic systems have a potential for maintaining a polymorphism in a large number of loci under selection without an input of new genetic variation.
Common terms used in genetics with multiple meanings are explained and a brief overview given of the four major areas of genetic epidemiology--the study of familial aggregation, segregation, cosegregation and association. Familial aggregation measures the potential for a trait to have a genetic aetiology. Segregation analysis uncovers single gene segregation. Cosegregation with genetic markers gives rise to linkage, which is used to locate trait genes on the genome. Association analysis is used for fine mapping, but rests on the assumption that linkage disequilibrium exists.
If a disease can be split into two or more groups on any criterion (clinical, biochemical, physiological or statistical) then the grouping can be tested to establish if genetically independent forms of the disease have been identified. The data required are simply the frequencies of the two disease groups in relatives of probands for each of the disease groups. A systematic search for such distinct groups is proposed in searches for genetic heterogeneity in familial diseases. In disease forms with overlapping, correlated genetic liabilities, the method of Falconer (1967) can be used to estimate the genetic correlation. However, when the groupings of the disease are confounded (such as one form precluding the other as in early and late onset diabetes) Falconer's method will be biased. Special methods of analysis to estimate the genetic parameters have been developed and are presented here. However, even when the groupings are confounded the Falconer method still gives reasonable estimates of the genetic correlation, in that they are unlikely to seriously mislead the investigator in the analysis and interpretation of observed data. In practice Falconer's simple method may be preferred to the more complex methods developed here because it involves fewer assumptions and can be applied over a wider range of circumstances.
Explore the source record for details and available documents.
A growing body of evidence suggests that genetic factors have an important influence on the onset and course of smoking. Here we review some of the statistical methods that have been used to test for genetic influences on smoking behaviour, with a particular focus on studies of large national twin samples. We show how many of the hypotheses that have been tested using a genetic model-fitting approach have also been reformulated using logistic regression models that will be more familiar to epidemiologists. Such an approach is more easily extended to allow for sociocultural, as well as genetic, influences on smoking behaviour. Using either approach, data are consistent in indicating that certainly in men, and possibly in women, genetic factors play an important role in predicting which individuals who become cigarette smokers progress to being long-term persistent smokers.
In analogy to the polygene determined morphological features, the DNA-fingerprint is also not suitable for statistical processing. Statements about the individuality are merely speculative. Frequencies of genes cannot be found, since it is impossible to determine which combinations of bands belong to one gene locus. Hence the DNA fingerprint enables the recognition of exclusions from paternity; it does not, however, allow a statistical analysis, no matter which method be employed.
Explore the source record for details and available documents.
Recent advances in molecular biology have made it possible for geneticists to collect data at the DNA sequence level, and so make observations on genes directly, instead of on characters affected by genes. This new nature of genetic data is posing new problems of statistical analyses, as is shown in this paper. After an explanation of the methodology of recombinant DNA studies, separate consideration is given to the statistical analyses of restriction-map and complete-sequence data. The initial observations from restriction-enzyme studies are of the number and sizes of the fragments produced by digesting a DNA region with the enzymes, and it is necessary to infer the location of the breakage points. Model-based analyses are then necessary to infer characteristics of the DNA sequence from the frequency of these break points. An important application of such restriction maps is the detection of human disease genes, with the most notable success to date being the chromosomal allocation of the gene for Huntington's chorea. Complete sequence data present conflicting problems of scale. The amount of information collected per individual is so large that analyses must always be done on a computer, but the number of individuals studied in a population is too low to allow the magnitudes of between- and within-population variation to be assessed. Statistical problems also include the detection and testing of sequence patterns and the comparison of sequences. Any analyses are going to have to incorporate the observed high levels of association between adjacent sequence members. While national data bases now contain about three million sequence elements, very little is known about the error rate in the generation or recording of these sequences.
We propose a model-driven approach for analyzing genomic expression data that permits genetic regulatory networks to be represented in a biologically interpretable computational form. Our models permit latent variables capturing unobserved factors, describe arbitrarily complex (more than pair-wise) relationships at varying levels of refinement, and can be scored rigorously against observational data. The models that we use are based on Bayesian networks and their extensions. As a demonstration of this approach, we utilize 52 genomes worth of Affymetrix GeneChip expression data to correctly differentiate between alternative hypotheses of the galactose regulatory network in S. cerevisiae. When we extend the graph semantics to permit annotated edges, we are able to score models describing relationships at a finer degree of specification.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
An approach to the study of the properties of genetic texts is proposed. It is based on the investigation of the frequencies of all possible words (subsequences) in a text. The most important effect is that the original text could be reconstructed completely without deletions and/or mistakes using the set of words which are met in the text as a single copy. The length of words for which the effect occurs is a measure of the text redundancy. Some real genetic sequences were studied as well.
Genetic studies of sexual isolation in Drosophila have generally failed to fully evaluate the effects of their sample size and recombination between markers on their conclusions. In this study we evaluate recombinational distances between markers in Drosophila pseudoobscura and D. persimilis, a species pair in which numerous genetic mapping studies have been performed. We conclude that, contrary to assertions, the inversions that distinguish these two species still allow for much recombination within most of their chromosome arms in F1 hybrid females. Such recombination may have caused previous mapping studies in these species to miss (or grossly underestimate) the effects of several genomic regions. We also evaluate the effects of sample size and recombination on genetic studies of sexual isolation in other Drosophila species groups. We conclude that some of these studies may have been heavily biased toward detecting only genes of large effect. Future studies of sexual isolation should be preceded by detailed statistical power analyses that determine the effects of recombination and sample size in the species pair being studied to avoid these complications.
A reduced representation model, which has been described in previous reports, was used to predict the folded structures of proteins from their primary sequences and random starting conformations. The molecular structure of each protein has been reduced to its backbone atoms (with ideal fixed bond lengths and valence angles) and each side chain approximated by a single virtual united-atom. The coordinate variables were the backbone dihedral angles phi and psi. A statistical potential function, which included local and nonlocal interactions and was computed from known protein structures, was used in the structure minimization. A novel approach, employing the concepts of genetic algorithms, has been developed to simultaneously optimize a population of conformations. With the information of primary sequence and the radius of gyration of the crystal structure only, and starting from randomly generated initial conformations, I have been able to fold melittin, a protein of 26 residues, with high computational convergence. The computed structures have a root mean square error of 1.66 A (distance matrix error = 0.99 A) on average to the crystal structure. Similar results for avian pancreatic polypeptide inhibitor, a protein of 36 residues, are obtained. Application of the method to apamin, an 18-residue polypeptide with two disulfide bonds, shows that it folds apamin to native-like conformations with the correct disulfide bonds formed.
Explore the source record for details and available documents.
Molecular markers that could stratify prostate cancer patients according to risk of disease progression would allow a significant improvement in the management of this clinically heterogeneous disease. In the present study, we analyzed the genetic profile of a consecutive series of 51 clinically confined prostate carcinomas and 27 benign prostatic hyperplasias using comparative genomic hybridization (CGH). We then added our findings to the existing literature data in order to perform a meta-analysis on a total of 294 prostate cancers with detailed CGH and clinicopathological information, using multivariate statistical methods that included principal component, hierarchical clustering, time of occurrence, and regression analyses. Whereas several genomic imbalances were shared by organ-confined, locally invasive, and metastatic prostate cancers, 6q and 10q losses and 7q and 8q gains were significantly more frequent in patients with extra-prostatic disease. Regression analysis indicated that 8q gain and 13q loss were the best predictors of locally invasive disease, whereas 8q gain and 6q and 10q losses were associated with metastatic disease. We propose a genetic pathway of prostate carcinogenesis with two distinct initiating events, namely, 8p and 13q losses. These primary imbalances are then preferentially followed by 8q gain and 6q, 16q, and 18q losses, which in turn are followed by a set of late events that make recurrent and metastatic prostate cancers genetically more complex. We conclude that significant differences exist in the genetic profile of organ-confined, locally invasive, and advanced prostate cancer and that genetic features may carry prognostic information independently of Gleason grade.