Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

Nursing process outcome linkage research: issues, current status, and health policy implications.

BACKGROUND: The use of large clinical datasets to assess the effectiveness of health care is of growing interest in continuing efforts to understand the impact of healthcare costs on quality. Correspondingly, there is a greater need to define and measure outcomes that are sensitive to nursing interventions. However, concerns exist about the ability to amass and use large clinical nursing datasets to assess the effectiveness of nursing interventions. Some nursing studies have used large clinical datasets to examine patterns of nursing diagnoses, interventions, and outcomes. Among patient populations, however, systematic effectiveness studies of nursing process and outcome linkages at the individual nurse and patient level of analysis are essentially nonexistent. This is largely the result of slow development of nursing classifications, reference terminologies, and reference information standards. Nursing information systems have an unprecedented potential for documentation of nursing practice, as well as the accumulation and analysis of large clinical datasets, to improve nursing performance, increase nursing knowledge, and provide data and information necessary for nursing to participate in the formulation of healthcare policy. OBJECTIVES: A literature search shows that a common framework is beginning to evolve that represents nursing's essential information, eg, the Nursing Minimum Data Set, Management Minimum Data Set, and several standardized nursing languages. Extensive research and other initiatives have produced 1) nursing languages and reference terminologies that span healthcare settings; 2) information models; and 3) standards for datasets supporting information systems. A number of issues remain, however, that concern the development of uniform nursing datasets, definitions of outcomes, quality of nursing data, information system design, and methods of data analysis. We review nursing process outcome research, clarify issues inherent in nursing effectiveness research, and discuss implications for nursing and health policy.

Health Policy↗

Evaluation of linked cancer registry and hospital records of breast cancer.

BACKGROUND: Information on the treatment of women with breast cancer in Australia is generally available only from special surveys. Analysis of routinely collected datasets may be more timely and cost effective, if the data are sufficiently accurate and complete. OBJECTIVE: To evaluate the accuracy and completeness of data on treatment in linked records of breast cancer from two routinely collected datasets. METHODS: The NSW Department of Health linked NSW Central Cancer Registry (CCR) records for 2,636 women diagnosed with breast cancer in NSW in 1992 to all hospital admission records in the NSW In-patient Statistics Collection (ISC) from January 1991 to June 1994. We queried the original paper records of subsets of women to identify missing or miscoded information and cases not notified to the CCR. We also compared the treatment data with data collected independently from the medical records of 19% of the women. RESULTS: ISC records linked to 89% of the CCR records. The CCR had identified 94.9% of women with breast cancer treated as hospital in-patients and 83% of these women had surgical treatment recorded in the ISC. The linked dataset under-estimated the percentage of women having breast-conserving therapy (-4%) and slightly over-estimated the percentage having mastectomy (+1%). We estimated that 42% of women treated surgically for breast cancer had actually had breast-conserving surgery, compared with 39% in the original dataset. There was no evident bias by age or by urban or rural residence in the under-recording of breast conservation. There was 94% agreement on the type of surgery between the linked dataset and the independent dataset.

Adult↗

Respiratory correlated cone beam CT.

A cone beam computed tomography (CBCT) scanner integrated with a linear accelerator is a powerful tool for image guided radiotherapy. Respiratory motion, however, induces artifacts in CBCT, while the respiratory correlated procedures, developed to reduce motion artifacts in axial and helical CT are not suitable for such CBCT scanners. We have developed an alternative respiratory correlated procedure for CBCT and evaluated its performance. This respiratory correlated CBCT procedure consists of retrospective sorting in projection space, yielding subsets of projections that each corresponds to a certain breathing phase. Subsequently, these subsets are reconstructed into a four-dimensional (4D) CBCT dataset. The breathing signal, required for respiratory correlation, was directly extracted from the 2D projection data, removing the need for an additional respiratory monitor system. Due to the reduced number of projections per phase, the contrast-to-noise ratio in a 4D scan reduced by a factor 2.6-3.7 compared to a 3D scan based on all projections. Projection data of a spherical phantom moving with a 3 and 5 s period with and without simulated breathing irregularities were acquired and reconstructed into 3D and 4D CBCT datasets. The positional deviations of the phantoms center of gravity between 4D CBCT and fluoroscopy were small: 0.13 +/- 0.09 mm for the regular motion and 0.39 +/- 0.24 mm for the irregular motion. Motion artifacts, clearly present in the 3D CBCT datasets, were substantially reduced in the 4D datasets, even in the presence of breathing irregularities, such that the shape of the moving structures could be identified more accurately. Moreover, the 4D CBCT dataset provided information on the 3D trajectory of the moving structures, absent in the 3D data. Considerable breathing irregularities, however, substantially reduces the image quality. Data presented for three different lung cancer patients were in line with the results obtained from the phantom study. In conclusion, we have successfully implemented a respiratory correlated CBCT procedure yielding a 4D dataset. With respiratory correlated CBCT on a linear accelerator, the mean position, trajectory, and shape of a moving tumor can be verified just prior to treatment. Such verification reduces respiration induced geometrical uncertainties, enabling safe delivery of 4D radiotherapy such as gated radiotherapy with small margins.

Algorithms↗

Comparison of normalization methods for CodeLink Bioarray data.

BACKGROUND: The quality of microarray data can seriously affect the accuracy of downstream analyses. In order to reduce variability and enhance signal reproducibility in these data, many normalization methods have been proposed and evaluated, most of which are for data obtained from cDNA microarrays and Affymetrix GeneChips. CodeLink Bioarrays are a newly emerged, single-color oligonucleotide microarray platform. To date, there are no reported studies that evaluate normalization methods for CodeLink Bioarrays. RESULTS: We compared five existing normalization approaches, in terms of both noise reduction and signal retention: Median (suggested by the manufacturer), CyclicLoess, Quantile, Iset, and Qspline. These methods were applied to two real datasets (a time course dataset and a lung disease-related dataset) generated by CodeLink Bioarrays and were assessed using multiple statistical significance tests. Compared to Median, CyclicLoess and Qspline exhibit a significant and the most consistent improvement in reduction of variability and retention of signal. CyclicLoess appears to retain more signal than Qspline. Quantile reduces more variability than Median in both datasets, yet fails to consistently retain more signal in the time course dataset. Iset does not improve over Median in either noise reduction or signal enhancement in the time course dataset. CONCLUSION: Median is insufficient either to reduce variability or to retain signal effectively for CodeLink Bioarray data. CyclicLoess is a more suitable approach for normalizing these data. CyclicLoess also seems to be the most effective method among the five different normalization strategies examined.

Algorithms↗

Differential prioritization between relevance and redundancy in correlation-based feature selection techniques for multiclass gene expression data.

BACKGROUND: Due to the large number of genes in a typical microarray dataset, feature selection looks set to play an important role in reducing noise and computational cost in gene expression-based tissue classification while improving accuracy at the same time. Surprisingly, this does not appear to be the case for all multiclass microarray datasets. The reason is that many feature selection techniques applied on microarray datasets are either rank-based and hence do not take into account correlations between genes, or are wrapper-based, which require high computational cost, and often yield difficult-to-reproduce results. In studies where correlations between genes are considered, attempts to establish the merit of the proposed techniques are hampered by evaluation procedures which are less than meticulous, resulting in overly optimistic estimates of accuracy. RESULTS: We present two realistically evaluated correlation-based feature selection techniques which incorporate, in addition to the two existing criteria involved in forming a predictor set (relevance and redundancy), a third criterion called the degree of differential prioritization (DDP). DDP functions as a parameter to strike the balance between relevance and redundancy, providing our techniques with the novel ability to differentially prioritize the optimization of relevance against redundancy (and vice versa). This ability proves useful in producing optimal classification accuracy while using reasonably small predictor set sizes for nine well-known multiclass microarray datasets. CONCLUSION: For multiclass microarray datasets, especially the GCM and NCI60 datasets, DDP enables our filter-based techniques to produce accuracies better than those reported in previous studies which employed similarly realistic evaluation procedures.

Animals↗

Comparison and evaluation of methods for generating differentially expressed gene lists from microarray data.

BACKGROUND: Numerous feature selection methods have been applied to the identification of differentially expressed genes in microarray data. These include simple fold change, classical t-statistic and moderated t-statistics. Even though these methods return gene lists that are often dissimilar, few direct comparisons of these exist. We present an empirical study in which we compare some of the most commonly used feature selection methods. We apply these to 9 publicly available datasets, and compare, both the gene lists produced and how these perform in class prediction of test datasets. RESULTS: In this study, we compared the efficiency of the feature selection methods; significance analysis of microarrays (SAM), analysis of variance (ANOVA), empirical bayes t-statistic, template matching, maxT, between group analysis (BGA), Area under the receiver operating characteristic (ROC) curve, the Welch t-statistic, fold change, rank products, and sets of randomly selected genes. In each case these methods were applied to 9 different binary (two class) microarray datasets. Firstly we found little agreement in gene lists produced by the different methods. Only 8 to 21% of genes were in common across all 10 feature selection methods. Secondly, we evaluated the class prediction efficiency of each gene list in training and test cross-validation using four supervised classifiers. CONCLUSION: We report that the choice of feature selection method, the number of genes in the genelist, the number of cases (samples) and the noise in the dataset, substantially influence classification success. Recommendations are made for choice of feature selection. Area under a ROC curve performed well with datasets that had low levels of noise and large sample size. Rank products performs well when datasets had low numbers of samples or high levels of noise. The Empirical bayes t-statistic performed well across a range of sample sizes.

Algorithms↗

Integrative missing value estimation for microarray data.

BACKGROUND: Missing value estimation is an important preprocessing step in microarray analysis. Although several methods have been developed to solve this problem, their performance is unsatisfactory for datasets with high rates of missing data, high measurement noise, or limited numbers of samples. In fact, more than 80% of the time-series datasets in Stanford Microarray Database contain less than eight samples. RESULTS: We present the integrative Missing Value Estimation method (iMISS) by incorporating information from multiple reference microarray datasets to improve missing value estimation. For each gene with missing data, we derive a consistent neighbor-gene list by taking reference data sets into consideration. To determine whether the given reference data sets are sufficiently informative for integration, we use a submatrix imputation approach. Our experiments showed that iMISS can significantly and consistently improve the accuracy of the state-of-the-art Local Least Square (LLS) imputation algorithm by up to 15% improvement in our benchmark tests. CONCLUSION: We demonstrated that the order-statistics-based integrative imputation algorithms can achieve significant improvements over the state-of-the-art missing value estimation approaches such as LLS and is especially good for imputing microarray datasets with a limited number of samples, high rates of missing data, or very noisy measurements. With the rapid accumulation of microarray datasets, the performance of our approach can be further improved by incorporating larger and more appropriate reference datasets.

Algorithms↗

Defining three dimensional chromatin structures of pediatric and adolescent B cells using primary B cell and EBV-immortalized B cell reference genomes.

BACKGROUND/PURPOSE: Knowledge of the 3D genome is essential to elucidate genetic mechanisms driving autoimmune diseases. The 3D genome is distinct for each cell type, and it is uncertain whether cell lines faithfully recapitulate the 3D architecture of primary human cells or whether developmental aspects of the pediatric immune system require use of pediatric samples. We undertook a systematic analysis of B cells and B cell lines to compare 3D genomic features encompassing risk loci for juvenile idiopathic arthritis (JIA), systemic lupus (SLE), and type 1 diabetes (T1D). METHODS: We isolated B cells from four healthy individuals, ages 9-17. HiChIP was performed using a CTCF antibody, and CTCF peaks were called within each sample separately. Peaks observed in all four samples were identified. CTCF loops were called within the pediatric samples using three CTCF peak datasets: 1) self-called CTCF consensus peaks called within the pediatric samples, 2) ENCODE's publicly available GM12878 CTCF ChIP-seq peaks, and 3) ENCODE's primary B cell CTCF ChIP-seq peaks from two adult females. Differential looping was assessed within the pediatric samples and each of the three peak datasets. RESULTS: The number of consensus peaks called in the pediatric samples was similar to that identified in ENCODE's GM12878 and primary B cell datasets. We observed&#x2009;<&#x2009;1% of loops that demonstrated significantly differential looping between peaks called within the pediatric samples themselves and when called using ENCODE GM12878 peaks. Significant looping differences were even fewer when comparing loops of the pediatric called peaks to those of the ENCODE primary B cell peaks. When querying loops found in juvenile idiopathic arthritis, type 1 diabetes, or systemic lupus erythematosus risk haplotypes, we observed significant differences in only 2.2%, 1.0%, and 1.3% loops, respectively, when comparing peaks called within the pediatric samples and ENCODE GM12878 dataset. The differences were even less apparent when comparing loops called with the pediatric vs ENCODE adult primary B cell peak datasets. CONCLUSION: The 3D chromatin architecture in B cells is similar across pediatric, adult, and EBV-transformed cell lines. This conservation of 3D structure includes regions encompassing autoimmune risk haplotypes. Thus, even for pediatric autoimmune diseases, publicly available adult B cell and cell line datasets may be sufficient for assessing effects exerted in the 3D genomic space.

Humans↗

Identification of cuproptosis-realated key genes and pathways in Parkinson's disease via bioinformatics analysis.

INTRODUCTION: Parkinson's disease (PD) is the second most common worldwide age-related neurodegenerative disorder without effective treatments. Cuproptosis is a newly proposed conception of cell death extensively studied in oncological diseases. Currently, whether cuproptosis contributes to PD remains largely unclear. METHODS: The dataset GSE22491 was studied as the training dataset, and GSE100054 was the validation dataset. According to the expression levels of cuproptosis-related genes (CRGs) and differentially expressed genes (DEGs) between PD patients and normal samples, we obtained the differentially expressed CRGs. The protein-protein interaction (PPI) network was achieved through the Search Tool for the Retrieval of Interacting Genes. Meanwhile, the disease-associated module genes were screened from the weighted gene co-expression network analysis (WGCNA). Afterward, the intersection genes of WGCNA and PPI were obtained and enriched using the Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG). Subsequently, the key genes were identified from the datasets. The receiver operating characteristic curves were plotted and a PPI network was constructed, and the PD-related miRNAs and key genes-related miRNAs were intersected and enriched. Finally, the 2 hub genes were verified via qRT-PCR in the cell model of the PD and the control group. RESULTS: 525 DEGs in the dataset GSE22491 were identified, including 128 upregulated genes and 397 downregulated genes. Based on the PPI network, 41 genes were obtained. Additionally, the dataset was integrated into 34 modules by WGCNA. 36 intersection genes found from WGCNA and PPI were significantly abundant in 7 pathways. The expression levels of the genes were validated, and 2 key genes were obtained, namely peptidase inhibitor 3 (PI3) and neuroserpin family I member 1 (SERPINI1). PD-related miRNAs and key genes-related miRNAs were intersected into 29 miRNAs including hsa-miR-30c-2-3p. At last, the qRT-PCR results of 2 hub genes showed that the expressions of mRNA were up-regulated in PD. CONCLUSION: Taken together, this study demonstrates the coordination of cuproptosis in PD. The key genes and miRNAs offer novel perspectives in the pathogenesis and molecular targeting treatment for PD.

Humans↗

Four-dimensional ultrasonography of the fetal heart using a novel Tomographic Ultrasound Imaging display.

OBJECTIVE: The objective of this study was to investigate the feasibility of examining the fetal heart with Tomographic Ultrasound Imaging (TUI) using four-dimensional (4D) volume datasets acquired with spatiotemporal image correlation (STIC). MATERIAL AND METHODS: One hundred and ninety-five fetuses underwent 4D ultrasonography (US) of the fetal heart with STIC. Volume datasets were acquired with B-mode (n=195) and color Doppler imaging (CDI) (n=168), and were reviewed offline using TUI, a new display modality that automatically slices 3D/4D volume datasets, providing simultaneous visualization of up to eight parallel planes in a single screen. Visualization rates for standard transverse planes used to examine the fetal heart were calculated and compared for volumes acquired with B-mode or CDI. Diagnoses by TUI were compared to postnatal diagnoses. RESULTS: (1) The four- and five-chamber views and the three-vessel and trachea view were visualized in 97.4% (190/195), 88.2% (172/195), and 79.5% (142/195), respectively, of the volume datasets acquired with B-mode; (2) these views were visualized in 98.2% (165/168), 97.0% (163/168), and 83.6% (145/168), respectively, of the volume datasets acquired with CDI; (3) CDI contributed additional diagnostic information to 12.5% (21/168), 14.2% (24/168) and 10.1% (17/168) of the four- and five-chamber and the three-vessel and trachea views; (4) cardiac anomalies other than isolated ventricular septal defects were identified by TUI in 16 of 195 fetuses (8.2%) and, among these, CDI provided additional diagnostic information in 5 (31.3%); (5) the sensitivity, specificity, positive- and negative-predictive values of TUI to diagnose congenital heart disease in cases where both B-mode and CDI volume datasets were acquired prenatally were 92.9%, 98.8%, 92.9% and 98.8%, respectively. CONCLUSION: Standard transverse planes commonly used to examine the fetal heart can be automatically displayed with TUI in the majority of fetuses undergoing 4D US with STIC. Due to the retrospective nature of this study, the results should be interpreted with caution and independently confirmed before this methodology is introduced into clinical practice.

Cardiac Volume↗

Comparative factor analysis models for an empirical study of EEG data, II: A data-guided resolution of the rotation indeterminacy.

In this paper (the second in a series), we consider a (generic) pair of datasets, which have been analyzed by the techniques of the previous paper. Thus, their "stable subspaces" have been established by comparative factor analysis. The pair of datasets must satisfy two confirmable conditions. The first is the "Inclusion Condition," which requires that the stable subspace of one of the datasets is nearly identical to a subspace of the other dataset's stable subspace. On the basis of that, we have assumed the pair to have similar generating signals, with stochastically independent generators. The second verifiable condition is that the (presumed same) generating signals have distinct ratios of variances for the two datasets. Under these conditions a small elaboration of some elementary linear algebra reduces the rotation problem to several eigenvalue-eigenvector problems. Finally, we emphasize that an analysis of each dataset by the method of Douglas and Rogers (1983) is an essential prerequisite for the useful application of the techniques in this paper. Nonempirical methods of estimating the number of factors simply will not suffice, as confirmed by simulations reported in the previous paper.

Animals↗

Using the STS and multinational cardiac surgical databases to establish risk-adjusted benchmarks for clinical outcomes.

One of the purposes of collecting data on cardiac surgical procedures, at a national level is to enable individual surgeons to improve quality and benchmark their own practice by making more accurate prospective prediction of outcome of each individual patient by using risk stratification based on previous local and national experiences. The past decade has seen a dramatic increase in the development of national cardiac surgical initiatives in many countries around the world. The size and extent of these databases has successfully allowed their use for patient risk stratification and preoperative risk modeling in four main aspects: patient selection and informed consent, coherent analysis of the determinants of patient outcomes, rationalizing unit management, and negotiations with external agencies. Approximately 610 cardiac surgical units presently contribute their patient data, containing pre-operative risk factors, to centralized national registries. There are currently nine different datasets used throughout the world to collect patient information. To harmonize the considerable diversity among these source materials, an International Dataset has been developed by a collaborative process among more than 50 cardiac surgeons around the world. Constructed around the Society of Thoracic Surgeons (STS) data format, the International Dataset brings in key elements from all the other datasets, allowing the sharing of data and cross-analysis, thus greatly expanding the pool of patients, and national sources, from which risk-stratifed outcomes can now be analyzed and unified. Unlike the STS dataset, the International Dataset incorporates EuroSCORE, a simple-to-use, validated patient risk stratification system, which has been rapidly adopted by large numbers of centers around the world for patient risk stratification, outcomes assessment, and improving patient informed consent. There are several benefits to collecting and centralizing national and international data: (1) understanding and defining basic demographics of patients undergoing cardiac surgery; (2) patient risk stratification and risk prediction at both a national and center-by-center level; (3) unit benchmarking, and development of effective nationally oriented and center-oriented quality improvement programs; (4) understanding and rationalizing resource utilization; and (5) use of data to leverage governments and other healthcare providers to affect policy. Cardiac surgical registries will soon attempt to track patients for longer follow-up periods after discharge in order to identify surgery-related deaths for more extended periods of time following surgery, thereby improving the monitoring and prediction of patient outcomes.

Benchmarking↗

[Scientific significance and prospective application of digitized virtual human].

As a cutting-edge research project, digitization of human anatomical information combines conventional medicine with information technology, computer technology, and virtual reality technology. Recent years have seen the establishment of, or the ongoing effort to establish various virtual human models in many countries, on the basis of continuous sections of human body that are digitized by means of computational medicine incorporating information technology to quantitatively simulate human physiological and pathological conditions, and to provide wide prospective applications in the fields of medicine and other disciplines. This article addresses 4 issues concerning the progress in virtual human model researches as the following: (1) Worldwide survey of sectioning and modeling of visible human. American visible human database was completed in 1994, which contains both a male and a female datasets, and has found wide application internationally. South Korea also finished the data collection for a male visible Korean human dataset in 2000. (2) Application of the dataset of Visible Human Project (VHP). This dataset has yielded plentiful fruits in medical education and clinical research, and further plans are proposed and practiced to construct a Physical Human and Physiological Human . (3) Scientific significance and prospect of virtual human studies. Digitized human dataset may eventually contribute to the development of many new high-tech industries. (4) Progress of virtual Chinese human project. The 174th session of Xiangshang Science Conferences held in 2001 marked the initiation of digitized virtual human project in China, and some key techniques have been explored. By now the data-collection process for 4 Chinese virtual human datasets have been successfully completed.

Anatomy, Cross-Sectional↗

Reshaping medical volumetric data for enhanced visualization.

Three dimensional volume datasets are now commonly produced in the medical sciences. These datasets are generated by observational equipment such as CT, MRI, and ultrasound. There is a significant amount of research in techniques to render these datasets quickly and more realistically. However, there is little or no work on intuitive methods to manipulate volume datasets and volume models. For example, one may want to view a colon stretched out or unraveled. While techniques exist to transform polygonal models, similar techniques are not available for volumetric data. In this work, we describe our methodology to "reshape volumes" and remap existing volumetric datasets using volumetric skeletons. We demonstrate our results by unraveling a 3D colon dataset and discuss the many potential uses of this new visualization methodology.

Anatomy, Cross-Sectional↗

Chinese visible human project.

Research on the digital visible human is of great significance and has considerable application value. The US visible human project created the first digital image dataset of a complete human (one male and one female) in 1995. To promote worldwide application-oriented visible human research, additional visible human datasets, representative of different populations of the world, are needed. The Chinese visible human (CVH) male (created in October 2002) and female (created in February 2003) Project achieved greater integrity of images, better blood vessel identification, and were free of organic disease. The most noteworthy technical advance of the Chinese visible human project (CVHP) was the construction of a low temperature laboratory, which prevented loss of small structures (including teeth, nasal conchae, and articular cartilage) from the milling surface. Thus, better integrity of images was achieved. To date, we have acquired five CVH datasets and volume rendered them for visualization on a PC. 3D reconstruction of some organs and structures has been completed and work to segment a complete dataset is under way. Although there is still a long way to go to make the visible human meet the application-oriented needs in various fields, progress is being made toward acquiring new datasets, performing segmentation, and setting up a platform of computer-assisted medicine. Here, we review the history and highlights of the CVHP and foresee its future development as well.

Anatomy↗

Do bilineal pedigrees represent a problem for linkage analysis? Basic principles and simulation results for single-gene diseases with no heterogeneity.

Some investigators have expressed concern--especially for psychiatric disorders--that bilineal pedigrees should not be included in linkage studies. This study compares the "informativeness" of bilineal and unilineal families for a homogeneous single-gene disorder. Three approaches were used: (1) simulation studies of three-generation pedigrees, (2) calculation of expected lod scores (ELODs) in nuclear families, and (3) calculation of Fisher's information number I(theta) in nuclear families. The simulation studies in (1) permitted a realistic comparison between bilineal datasets and purely unilineal ones. The calculations in nuclear families in (2) and (3) then made it possible to analyze the sources of information loss in bilineal families. Overall, in datasets of five three-generation pedigrees each, the drop in mean maximum lod score was approximately 50% from purely unilineal datasets to extremely bilineal ones. In less-extreme bilineal datasets, which are closer to most real data than the extremely bilineal ones, the drops in lod score were very small--less than 10% in some, and practically zero in others. The details will vary, depending on size and structure of the pedigree, genetic model, true value of the recombination fraction, and informativeness of the marker. However, these results imply that the information loss due to bilineality is not necessarily very great. The nuclear-family calculations showed that for phase-known matings there is relatively little information loss in bilineal families, but for phase-unknown matings there the loss is much greater. In conclusion, for single-gene disorders with no genetic heterogeneity, whereas bilineal families can be less informative than comparable unilineal families, they are not so much less informative that they should automatically be discarded from linkage datasets. The implications of bilineal pedigrees for linkage studies of heterogeneous disorders are also discussed.

Computer Simulation↗

Mining the structural knowledge of high-dimensional medical data using isomap.

The paper describes an application of a new, non-linear dimensionality reduction method, named Isomap, for mining the structural knowledge from high-dimensional medical data. The algorithm was evaluated on two publicly available medical datasets: the pathological dataset of breast cancer (241 malignant samples) and the gene expression dataset from the lung (186 tumours). It was found by Isomap that the approximate intrinsic dimensionalities of these two datasets were as low as three. The spatial structures of both datasets were presented in low-dimensional space. Isomap, as a general tool for dimensionality reduction analysis, is helpful in revealing the nonlinear structural knowledge of high-dimensional medical data.

Algorithms↗

[3-dimensional echo-enhanced transcranial Doppler ultrasound diagnosis].

Echo-enhancing agents improve the signal intensity of transcranial Doppler signals, enabling a novel approach of three-dimensional transcranial vascular imaging by Doppler ultrasound. The basic principle, system requirements and early clinical results with a custom built system are described. Transcranial color Doppler imaging was performed through the temporal bone acoustic window. During i.v. administration of the transpulmonary stable, galactose-based echo-enhancer Levovist (Schering) the video output of the ultrasound scanner was digitized and the spatial position of the recorded frames was simultaneously registered using a mechanical position sensor. After automatic segmentation of the color information, the 3D datasets were reconstructed offline using a Unix-based workstation (Silicon Graphics). Visualization was achieved by maximum intensity projection or surface visualization techniques. Administration of Levovist resulted in good enhancement of the vascular Doppler signal intensity, enabling acquisition of a 3D dataset of the complete circle of Willis with an imaging window of approximately 3-5 min for one i.v. injection. The vascularity of tumors could be recorded as a 3D dataset and further analyzed. The power Doppler technique with echoenhancement proved a valuable tool for 3D dataset recording. 3D datasets clearly facilitated the diagnosis of the vascular anatomy and lesion vascularity and provided additional information on localization of feeders, vascular displacement and extent of tumor vascularity.

Blood Flow Velocity↗