Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

Assessment of methods for amino acid matrix selection and their use on empirical data shows that ad hoc assumptions for choice of matrix are not justified.

BACKGROUND: In recent years, model based approaches such as maximum likelihood have become the methods of choice for constructing phylogenies. A number of authors have shown the importance of using adequate substitution models in order to produce accurate phylogenies. In the past, many empirical models of amino acid substitution have been derived using a variety of different methods and protein datasets. These matrices are normally used as surrogates, rather than deriving the maximum likelihood model from the dataset being examined. With few exceptions, selection between alternative matrices has been carried out in an ad hoc manner. RESULTS: We start by highlighting the potential dangers of arbitrarily choosing protein models by demonstrating an empirical example where a single alignment can produce two topologically different and strongly supported phylogenies using two different arbitrarily-chosen amino acid substitution models. We demonstrate that in simple simulations, statistical methods of model selection are indeed robust and likely to be useful for protein model selection. We have investigated patterns of amino acid substitution among homologous sequences from the three Domains of life and our results show that no single amino acid matrix is optimal for any of the datasets. Perhaps most interestingly, we demonstrate that for two large datasets derived from the proteobacteria and archaea, one of the most favored models in both datasets is a model that was originally derived from retroviral Pol proteins. CONCLUSION: This demonstrates that choosing protein models based on their source or method of construction may not be appropriate.

Amino Acid Substitution↗

Will the real disease gene please stand up?

A common dilemma arising in linkage studies of complex genetic diseases is the selection of positive signals, their follow-up with association studies and discrimination between true and false positive results. Several strategies for overcoming these issues have been devised. Using the Genetic Analysis Workshop 14 simulated dataset, we aimed to apply different analytical approaches and evaluate their performance in discerning real associations. We considered a) haplotype analyses, b) different methods adjusting for multiple testing, c) replication in a second dataset, and d) exhaustive genotyping of all markers in a sufficiently powered, large sample group. We found that haplotype-based analyses did not substantially improve over single-point analysis, although this may reflect the low levels of linkage disequilibrium simulated in the datasets provided. Multiple testing correction methods were in general found to be over-conservative. Replication of nominally positive results in a second dataset appears to be less stringent, resulting in the follow-up of false positives. Performing a comprehensive assay of all markers in a large, well-powered dataset appears to be the most effective strategy for complex disease gene identification.

Chromosome Mapping↗

Accuracy and completeness of mortality data in the Department of Veterans Affairs.

BACKGROUND: One of the national mortality databases in the U.S. is the Beneficiary Identification and Record Locator Subsystem (BIRLS) Death File that contains death dates of those who have received any benefits from the Department of Veterans Affairs (VA). The completeness of this database was shown to vary widely from cohort to cohort in previous studies. Three other sources of death dates are available in the VA that can complement the BIRLS Death File. The objective of this study is to evaluate the completeness and accuracy of death dates in the four sources available in the VA and to examine whether these four sources can be combined into a database with improved completeness and accuracy. METHODS: A random sample of 3,000 was drawn from 8.3 million veterans who received benefits from the VA between 1997 and 1999 and were alive on January 1, 1999 according to at least one source. Death dates found in BIRLS Death File, Medical SAS Inpatient Datasets, Medicare Vital Status, and Social Security Administration (SSA) Death Master File were compared with dates obtained from the National Death Index. A combined dataset from these sources was also compared with National Death Index dates. RESULTS: Compared with the National Death Index, sensitivity (or the percentage of death dates correctly recorded in a source) was 77.4% for BIRLS Death File, 12.0% for Medical SAS Inpatient Datasets, 83.2% for Medicare Vital Status, and 92.1% for SSA Death Master File. Over 95% of death dates in these sources agreed exactly with dates from the National Death Index. Death dates in the combined dataset demonstrated 98.3% sensitivity and 97.6% exact agreement with dates from the National Death Index. CONCLUSION: The BIRLS Death File is not an adequate source of mortality data for the VA population due to incompleteness. When the four sources of mortality data are carefully combined, the resulting dataset can provide more timely data for death ascertainment than the National Death Index and has comparable accuracy and completeness.

Journal Article↗

Preoperative calculation of risk for prolonged intensive care unit stay following coronary artery bypass grafting.

OBJECTIVE: Patients who have prolonged stay in intensive care unit (ICU) are associated with adverse outcomes. Such patients have cost implications and can lead to shortage of ICU beds. We aimed to develop a preoperative risk prediction tool for prolonged ICU stay following coronary artery surgery (CABG). METHODS: 5,186 patients who underwent CABG between 1st April 1997 and 31st March 2002 were analysed in a development dataset. Logistic regression was used with forward stepwise technique to identify preoperative risk factors for prolonged ICU stay; defined as patients staying longer than 3 days on ICU. Variables examined included presentation history, co-morbidities, catheter and demographic details. The use of cardiopulmonary bypass (CPB) was also recorded. The prediction tool was tested on validation dataset (1197 CABG patients between 1st April 2003 and 31st March 2004). The area under the receiver operating characteristic (ROC) curve was calculated to assess the performance of the prediction tool. RESULTS: 475 (9.2%) patients had a prolonged ICU stay in the development dataset. Variables identified as risk factors for a prolonged ICU stay included renal dysfunction, unstable angina, poor ejection fraction, peripheral vascular disease, obesity, increasing age, smoking, diabetes, priority, hypercholesterolaemia, hypertension, and use of CPB. In the validation dataset, 8.1% patients had a prolonged ICU stay compared to 8.7% expected. The ROC curve for the development and validation datasets was 0.72 and 0.74 respectively. CONCLUSION: A prediction tool has been developed which is reliable and valid. The tool is being piloted at our institution to aid resource management.

Aged↗

Fine mapping of genes within the IDDM8 region in rheumatoid arthritis.

The IDDM8 region on chromosome 6q27, first identified as a susceptibility locus for type 1 diabetes, has previously been linked and associated with rheumatoid arthritis (RA). The region contains a number of potential candidate genes, including programmed cell death 2 (PDCD2), the proteosome subunit beta type 1 (PSMB1), delta-like ligand 1 (DLL-1) and TATA box-binding protein (TBP) amongst others. The aim of this study was to fine map the IDDM8 region on chromosome 6q27, focusing on the genes in the region, to identify polymorphisms that may contribute to susceptibility to RA and potentially to other autoimmune diseases. Validated single nucleotide polymorphisms (SNPs; n = 65) were selected from public databases from the 330 kb region of IDDM8. These were genotyped using Sequenom MassArray genotyping technology in two datasets; the test dataset comprised 180 RA cases and 180 controls. We tested 50 SNPs for association with RA and any significant associations were genotyped in a second dataset of 174 RA cases and 192 controls, and the datasets were combined before analysis. Association analysis was performed by chi-square test implemented in Stata software and linkage disequilibrium and haplotype analysis was performed using Helix tree version 4.1. There was initial weak evidence of association, with RA, of a number of SNPs around the loc154449 putative gene and within the KIAA1838 gene; however, these associations were not significant in the combined dataset. Our study has failed to detect evidence of association with any of the known genes mapping to the IDDM8 locus with RA.

Adult↗

A multi-ancestry polygenic risk score for body mass index predicts longitudinal weight change.

BACKGROUND: Identifying individuals at risk for future weight gain is challenging, partly because associations with traditional clinical risk factors may be biased by confounding and reverse causation. Polygenic risk scores (PRS) provide a stable, lifelong measure of genetic predisposition to obesity. However, existing PRS have not been evaluated for their association with longitudinal weight change in adulthood and often lack generalizability across diverse genetic ancestry groups. METHODS: We conducted ancestry-specific genome-wide association study meta-analyses of body mass index (BMI) in populations of European, African or African American, Admixed American, East Asian, and South Asian ancestries and developed ancestry-specific PRS. A multi-ancestry polygenic risk score (MAPRS) was trained using ancestry-specific PRS in a model selection dataset (N = 39,685) from the All of Us Research Program (AoU). We evaluated the MAPRS in an independent AoU model evaluation dataset (N = 158,743) for BMI prediction and in a separate AoU test dataset (N = 78,219) with repeated measurements over 1.5-2.5 years for weight change prediction. The outcomes included change in BMI and ≥ 10% or ≥ 5% total body weight (TBW) gain. We further examined the relationship between MAPRS and 12 clinical risk factors commonly comorbid with obesity in relation to weight change. RESULTS: The MAPRS captured 7.05% of the variance in measured BMI in the AoU model evaluation dataset and demonstrated improved generalizability across all non-European genetic ancestry groups. In the AoU test dataset, conditioned on baseline BMI at the second-to-last measurement, a one SD increase in MAPRS was associated with a 0.16 kg/m2 increase in future BMI (standard error = 0.012 kg/m2; p-value = 2.2 × 10-39), 1.27-fold increased odds of experiencing ≥ 10% TBW gain (95% CI: 1.24-1.31; p-value = 1.4 × 10-55), and 1.15-fold increased odds of experiencing ≥ 5% TBW gain (95% CI: 1.13-1.18; p-value = 2.8 × 10-39). These associations were observed across all genetic ancestry groups and remained highly consistent after adjustment for any clinical risk factor. In contrast, most clinical risk factors demonstrated inconsistent or weaker associations with weight change outcomes. CONCLUSIONS: We developed an MAPRS for BMI that represents a robust and generalizable risk factor for longitudinal weight gain in adulthood, providing a foundation for genetically informed risk stratification and earlier, more targeted obesity prevention strategies.

Humans↗

Building an anonymized catalogued radiology museum in PACS: a feasibility study.

The aim of this study was to test the feasibility of a software application that would allow the anonymization and cataloguing of whole DICOM datasets in order to build searchable radiology museums within PACS. The application was developed on a dedicated networked PC, using C# and HL7 coding. Whole DICOM datasets were pushed from PACS to a networked PC on which the application, Museum Builder, was developed. Museum Builder works by replacing the patient specific data (the forename, surname and hospital number) within each header of each DICOM file with terms from anatomical and surgical sieve menus. The date of birth is anonymized to 1 January of the same year. Whole DICOM datasets comprising hundreds of images can be anonymized and catalogued in a single episode. Museum Builder primes PACS with an HL7 script to receive a "new" patient. DICOM datasets are then pushed back to PACS where they are added to the database as "new" cases. The museum cases can then be searched for, on PACS, by any combination of terms that correspond to appropriate anatomical units, surgical sieve headings or radiological specialty. New radiology reports containing clinical histories, radiological descriptions, differential diagnoses and discussion can be added through the report window. Our institution has developed and used this tool to generate a PACS based radiology museum containing not only full DICOM datasets, but also relevant histological and clinical photographs. In conclusion, this technique offers a mechanism for generating anonymized catalogued radiology museums in PACS. Museum Builder represents a working prototype that demonstrates some of the archiving functions that are expected by teaching institutions from PACS.

Cataloging↗

Three-dimensional MRI of the male urethrae with implanted artificial sphincters: initial results.

The aim of this study was to develop a method for simultaneous 3D visualization of a new type of artificial urethral sphincter (AUS) and adjacent urinary structures. Serial MR tomograms were acquired from seven men after AUS implantation. 3D reconstruction was performed by thresholding original (positive) and inverted (negative) image intensity and by subsequently fusing positive and negative images. Results show that the bladder, cuff and balloons of the AUS of originally high intensity were imaged in 3D by thresholding the positive datasets. The urethrae and corpora cavernosa penis of originally low intensity were displayed in 3D by thresholding the negative datasets. Fusion of the positive and negative datasets allowed simultaneous visualization of the AUS complex and adjacent urinary structures. All the structures of interest were also clearly seen by interactive multiplanar reformatting. Coronal tomographic datasets provided better 3D and reformatted 2D images than sagittal and transverse datasets. This technique offers a simple means for evaluating the complex urethral anatomy and the AUS, and has potential for improved 3D visualization of many other complex morphological and pathological conditions.

Aged↗

Identification of differentially expressed genes in multiple microarray experiments using discrete fourier transform.

Research in the post-genome sequence era has been shifting towards a functional understanding of the roles and relationships between different genes in different conditions. While the advances in genetic expression profiling techniques including microarrays enable detailed and genome-scale measurements, the extraction of meaningful information from large datasets remains a challenging task. Here, we propose a novel method of generating gene differential expression profiles such that gene expression values from one dataset can be directly compared with those of another dataset. A simplified Discrete Fourier Transform is applied to interposed gene expression values, thereby generating the 'spectra' for a pair of conditions. Using this technique, differentially expressed genes produce higher amplitudes at the Nyquist Frequency. By measuring the phase of the 'spectra' generated, the over- and under-expressed nature of the genes can be identified. This method was validated using two sets of GeneChip array data, one from prostate cancer related dataset and the other from macular degeneration related dataset. The genes identified as differentially expressed by our method were found to be similar to those published using their preferred methods. Based on our findings, the proposed DFT method could be used efficiently in identifying differentially expressed genes from multiple-array experiments from two different conditions.

Fourier Analysis↗

C-arm-based three-dimensional navigation: a preliminary feasibility study.

OBJECTIVE: With the new Siremobil Iso-C3D C-arm, three-dimensional (3D) datasets can be acquired intraoperatively in near-real time. Preliminary studies investigated the advantages of this system for depiction in joint and spinal surgery. Three-dimensional navigation seems feasible using the DICOM dataset of the Siremobil Iso-C3D in navigation devices. An experimental study was designed to investigate the feasibility and accuracy of this new technique. MATERIALS AND METHODS: After implantation of fiducial markers (titanium mini-screws, Leibinger), a Siremobil Iso-C3D C-arm with standard imaging options was used to acquire pre-interventional 3D datasets of the specimens. These isotropic voxel data were transferred via DICOM to a medivision navigation system using the spine module. After registration of the fiducials, a total of 20 pedicle screws were implanted (in 4 artificial-bone vertebral bodies and 6 cadaver vertebrae in situ) with the use of the navigation system in real-time mode. Post-interventionally, Iso-C3D and CT scans were obtained to control for implant position in the cadaver study. RESULTS: Fiducial marker implantation and registration require a special protocol to ensure correct identification and patient orientation in the DICOM dataset. The obtained accuracy was within 2 mm. Post-interventional imaging of the cadaveric vertebrae showed 10 of 12 screws to be correctly placed, with the other two in marginal intraosseous positions. CONCLUSIONS: Three-dimensional navigation with the Siremobil Iso-C3D data set is feasible, the accuracy being comparable to that of CT-based navigation and adequate for clinical interventions. Fiducial marker-based registration allows navigation of different bones in the same dataset without additional 3D scanning. This method is very useful as an additional tool in registration-free Iso-C3D-based navigation, since the navigation system allows the use of only one dynamic reference base (DRB).

Bone Screws↗

Comparison of mutual information-based warping accuracy for fusing body CT and PET by 2 methods: CT mapped onto PET emission scan versus CT mapped onto PET transmission scan.

UNLABELLED: This article assesses the resulting accuracies of 2 registration methods using the same multimodal mutual information registration algorithm. In the indirect, fusion method, the CT dataset is warped onto the PET transmission scan, and then the patient's attenuation-corrected emission dataset is substituted for the transmission dataset. In the direct, fusion method, the CT is warped directly onto the attenuation-corrected emission dataset. METHODS: CT and (18)F-FDG PET image datasets from 14 subjects with malignant lesions in the thorax were registered. In both CT and PET imaging acquisitions, the patient's arms were at the patient's side, resting on the scanning couch in a manner similar to that of routine PET acquisition procedures. The accuracy of the 2 warping registrations was assessed by measuring the distance between lesion centroids on CT and PET emission after fusion. RESULTS: The indirect method has a statistically smaller mean error, 6.2 mm, than the direct method, 10.6 mm. CONCLUSION: The indirect method appears to be the more accurate/reliable choice for fusing body CT and FDG PET.

Algorithms↗

[Current problems in the data acquisition of digitized virtual human and the countermeasures].

As a relatively new field of medical science research that has attracted the attention from worldwide researchers, study of digitized virtual human still awaits long-term dedicated effort for its full development. In the full array of research projects of the integrated Virtual Chinese Human project, virtual visible human, virtual physical human, virtual physiome, and intellectualized virtual human must be included as the four essential constitutional opponents. The primary importance should be given to solving the problems concerning the data acquisition for the dataset of this immense project. Currently 9 virtual human datasets have been established worldwide, which are subjected to critical analyses in the paper with special attention given to the problems in the data storage and the techniques employed, for instance, in these datasets. On the basis of current research status of Virtual Chinese Human project, the authors propose some countermeasures for solving the problems in the data acquisition for the dataset, which include (1) giving the priority to the quality control instead of merely racing for quantity and speed, and (2) improving the setting up of the markers specific for the tissues and organs to meet the requirement from information technology, (3) with also attention to the development potential of the dataset which should have explicit pertinence to specific actual applications.

Anatomy, Cross-Sectional↗

Internal correspondence analysis of codon and amino-acid usage in thermophilic bacteria.

Starting from two datasets of codon usage in coding sequences from mesophilic and thermophilic bacteria, we used internal correspondence analysis to study the variability of codon usage within and between species, and within and between amino acids. The first dataset included 18,958,458 codons from 58,482 coding sequences from completely sequenced genomes of 25 species, along with 6,793,581 dinucleotides from 21,876 intergenic spaces. The second dataset, with partially sequenced genomes, included 97,095,873 codons from 293 bacterial species. Results were consistent between the two datasets. The trend for the amino-acid composition of thermophilic proteins was found to be under the control of a pressure at the nucleic acid level, not a selection at the protein level. This effect was not present in intergenic spaces, ruling out a pressure at the DNA level. The pattern at the mRNA level was more complex than a simple purine enrichment of the sense strand of coding sequences. Outliers in the partial genome dataset introduced a note of caution about the interpretation of temperature as the direct determinant of the trend observed in thermophiles. The surprising lack of selection on the amino-acid content of thermophilic proteins suggests that the amino-acid repertoire was set up in a hot environment.

Bacteria↗

Empirical analysis of the STR profiles resulting from conceptual mixtures.

Samples containing DNA from two or more individuals can be difficult to interpret. Even ascertaining the number of contributors can be challenging and associated uncertainties can have dramatic effects on the interpretation of testing results. Using an FBI genotypes dataset, containing complete genotype information from the 13 Combined DNA Index System (CODIS) loci for 959 individuals, all possible mixtures of three individuals were exhaustively and empirically computed. Allele sharing between pairs of individuals in the original dataset, a randomized dataset and datasets of generated cousins and siblings was evaluated as were the number of loci that were necessary to reliably deduce the number of contributors present in simulated mixtures of four or less contributors. The relatively small number of alleles detectable at most CODIS loci and the fact that some alleles are likely to be shared between individuals within a population can make the maximum number of different alleles observed at any tested loci an unreliable indicator of the maximum number of contributors to a mixed DNA sample. This analysis does not use other data available from the electropherograms (such as peak height or peak area) to estimate the number of contributors to each mixture. As a result, the study represents a worst case analysis of mixture characterization. Within this dataset, approximately 3% of three-person mixtures would be mischaracterized as two-person mixtures and more than 70% of four-person mixtures would be mischaracterized as two- or three-person mixtures using only the maximum number of alleles observed at any tested locus.

Alleles↗

Scalable data servers for large multivariate volume visualization.

Volumetric datasets with multiple variables on each voxel over multiple time steps are often complex, especially when considering the exponentially large attribute space formed by the variables in combination with the spatial and temporal dimensions. It is intuitive, practical, and thus often desirable, to interactively select a subset of the data from within that high-dimensional value space for efficient visualization. This approach is straightforward to implement if the dataset is small enough to be stored entirely in-core. However, to handle datasets sized at hundreds of gigabytes and beyond, this simplistic approach becomes infeasible and thus, more sophisticated solutions are needed. In this work, we developed a system that supports efficient visualization of an arbitrary subset, selected by range-queries, of a large multivariate time-varying dataset. By employing specialized data structures and schemes of data distribution, our system can leverage a large number of networked computers as parallel data servers, and guarantees a near optimal load-balance. We demonstrate our system of scalable data servers using two large time-varying simulation datasets.

Computer Graphics↗

A rapid and accurate method to realign PET scans utilizing image edge information.

UNLABELLED: Movement during or between PET examinations is a common and serious problem. Consequently, there is a great need for rapid, accurate and robust methods to realign image sets. METHODS: Derivative information from the image sets was used to extract areas containing edge information. Image similarity between a reference dataset and a misaligned dataset was evaluated for these areas. Powell's method for function minimization was used to find the set of translations and rotations along and around the axes that maximized image similarity. The method was validated by realigning image sets with a known misalignment. Image sets used for validation included brain studies using several different tracers and heart studies using labeled acetate or water. RESULTS: The method was capable of labeled acetate or water. RESULTS: The method was capable of realigning brain datasets using the same tracer with an accuracy of 0.2 mm and 0.2 degrees along and around all axes. The same accuracy was obtained for datasets with as few as a total of 800,000 counts. Brain studies utilizing different tracers with markedly dissimilar regional uptake patterns were realigned with an accuracy of 1.5 mm and 1.5 degrees. Heart studies using water or acetate were realigned with an accuracy of 0.2 mm and 0.4 degrees along and around all axes. Realignment of a heart study containing a large focal uptake defect against a dataset without defect produced errors no greater than 1.0 mm and 1.0 degree. CONCLUSION: The use of derivative information provides a useful method to accurately realign PET image sets. It is rapid and noise-insensitive enough to allow for its routine use in dynamic studies.

Acetates↗

Frequency difference gating: a multivariate method for identifying subsets that differ between samples.

BACKGROUND: In multivariate distributions (for example, in 3- or more color flow cytometric datasets), it can become difficult or impossible to identify populations that differ between samples based only on a combination of univariate or bivariate displays. Indeed, it is possible that such differences can only be identified in "n"-dimensional space, where "n" is the number of parameters measured. Therefore, computer assisted identification of such differences is necessary. Such a method could be used to identify responses (i.e., by comparing cell samples before and after stimulation) in exquisite detail by allowing complete analysis of the collected data on only those events which have responded. METHODS: Multivariate Probability Binning can be used to compare different datasets to identify the distance and statistical significance of a difference between the distributions. An intermediate step in the algorithm provides access to the actual locations within the n-dimensional comparison which are most different between the distributions. Gates based on collections of hyper-rectangular bins can then be applied to datasets, thereby selecting those events (or clusters of events) that are different between samples. We term this process Frequency Difference Gating. RESULTS: Frequency Difference Gating was used in several test scenarios to evaluate its utility. First, we compared PBMC subsets identified by solely by immunofluorescence staining: based on this training data set, the algorithm automatically generated an accurate forward and side-scatter gate to identify lymphocytes. Second, we applied the algorithm to identify subtle differences between CD4 memory subsets based on 8-color immunophenotyping data. The resulting 3-dimensional gate could resolve cells subsets much more frequent in one subset compared to the other; no combination of two-dimensional gates could accomplish this resolution. Finally, we used the algorithm to compare B cell populations derived from mice of different ages or strains, and found that the algorithm could find very subtle differences between the populations. CONCLUSION: Frequency Difference Gating is a powerful tool that automates the process of identifying events comprising underlying differences between samples. It is not a clustering tool; it is not meant to identify subsets in multidimensional space. Importantly, this method may reveal subtle changes in small populations of cells, changes that only occur simultaneously in multiple dimensions in such a way that identification by univariate or bivariate analyses is impossible. Finally, the method may significantly aid in the analysis of high-order multivariate data (i.e., 6-12 color flow cytometric analyses), where identification of differences between datasets becomes so time-consuming as to be impractical. Published 2001 Wiley-Liss, Inc.

Algorithms↗

The heterogeneity problem. I: Separating genetic from environmental forms of the same disease.

The purposes of this work were 1) to reparameterize the likelihood used in segregation analysis in a way particularly suited to detecting heterogeneity (the result of the analysis is a parameter giving the proportion of families with the genetic form of the disease in the dataset) and 2) to test how well this reparameterization works using simulation. We assume that a dataset contains nuclear family data, with some of the families having a form of the disease that is environmentally caused and the others with a genetic form of the disease. In this study, we considered the case where the genetic form is a simple recessive and the environmental form a random model. The underlying parameters were the gene frequency, q, and the frequency of sporadics, R. We reparameterized the likelihood in terms of alpha, the percentage of genetic families in the dataset, which we attempt to estimate. We contrast the estimates of alpha with the population heterogeneity as reflected in the estimates of q and R. For the simulation, nuclear families are generated. Genetic families were simulated with a mendelian recessive pattern and environmental families according to a simple random model. Over a wide range of generating parameters, estimates of alpha were good, differing from the "true" values by only a few percent. Estimates of q and R, on the other hand, ranged from fair to poor. Our results indicate that the amount of heterogeneity in a dataset can be accurately estimated using segregation analysis, even when estimates of the gene frequency and penetrance among sporadics are unreliable.

Environment↗