Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,189 records · Page 66Linked to original sources

Reducing false positives in molecular pattern recognition.

In the search for new cancer subtypes by gene expression profiling, it is essential to avoid misclassifying samples of unknown subtypes as known ones. In this paper, we evaluated the false positive error rates of several classification algorithms through a 'null test' by presenting classifiers a large collection of independent samples that do not belong to any of the tumor types in the training dataset. The benchmark dataset is available at www2.genome.rcast.u-tokyo.ac.jp/pm/. We found that k-nearest neighbor (KNN) and support vector machine (SVM) have very high false positive error rates when fewer genes (<100) are used in prediction. The error rate can be partially reduced by including more genes. On the other hand, prototype matching (PM) method has a much lower false positive error rate. Such robustness can be achieved without loss of sensitivity by introducing suitable measures of prediction confidence. We also proposed a cluster-and-select technique to select genes for classification. The nonparametric Kruskal-Wallis H test is employed to select genes differentially expressed in multiple tumor types. To reduce the redundancy, we then divided these genes into clusters with similar expression patterns and selected a given number of genes from each cluster. The reliability of the new algorithm is tested on three public datasets.

Computational Biology↗

Documentation of family violence in New Zealand general practice.

AIM: To determine the rate of family violence documented during general practice consultations and describe clinical presentations. METHOD: A dataset of 447,809 computerised consultations involving 143,634 patients from 41 general practices throughout New Zealand was examined to identify consultations recording family violence issues. The documentation rate was determined and the subset analysed. RESULTS: A subset of 337 consultations from the 447,809 examined (0.075%) involved a family violence issue. This subset included 311 patients, 0.2% of the 143,634 patients in the original 6-month dataset. 225 (81%) of the patients in the subset were female. Family violence was the main reason for presentation in 137 (40%) consultations. The perpetrator was identified as the partner in 134 (40%) consultations, as the parent in 54 (16%) consultations and the patient identified themselves as the possible abuser in 17 (5%) consultations. Physical abuse (42%) and sexual abuse (26%) was most commonly mentioned. Past abuse (42%) was discussed as often as current abuse (41%). Depression and anxiety disorders were documented in 59 (18%) of these consultations. CONCLUSIONS: The number of consultations documenting family violence is low in this dataset. Such information is not always recorded, however GPs can also be reluctant to ask about, and patients can be hesitant to disclose family violence issues. The number of consultations involving the perpetrator was higher than expected. GPs require training to deal with both the victim and the perpetrator of family violence.

Documentation↗

Evaluation of methods to detect interhemispheric asymmetry on cerebral perfusion SPECT: application to epilepsy.

UNLABELLED: Detecting perfusion interhemispheric asymmetry in neurologic nuclear medicine imaging is an interesting approach to epilepsy. METHODS: This study compared 4 methods that detect interhemispheric asymmetries of brain perfusion in SPECT. The first (M1) was conventional side-by-side expert-based visual interpretation of SPECT. The second (M2) was visual interpretation assisted by an interhemispheric difference (IHD) volume. The last 2 were automatic methods: unsupervised analysis using volumes of interest (M3) and unsupervised analysis of the IHD volume (M4). Use of these methods to detect possible perfusion asymmetry was compared on 60 simulated SPECT datasets by controlling the presence and location of asymmetries. From the detection results, localization receiver operating characteristic curves were generated and areas under curves were estimated and compared. Finally, the methods were applied to analyze interictal SPECT datasets to localize the epileptogenic focus in temporal lobe epilepsies. RESULTS: This study showed an improvement in asymmetry detection on SPECT images with the methods using IHD volume (M2 and M4), in comparison with the other methods (M1 and M3). However, the most useful method for analyzing clinical SPECT datasets appeared to be visual inspection assisted by the IHD volume, since the automatic method using the IHD volume was less specific. CONCLUSION: The use of quantitative methods can improve performance in detection of perfusion asymmetry over visual inspection alone.

Adolescent↗

Evaluation of sequest result filter-Xcorr and Unified Score.

OBJECTIVE: To estimate the effect of two simple filters, two or more positive peptide filter and Unified Score filter on the true positive rate of protein and peptide. METHODS: Twenty-two LC-MS/MS datasets were from 18 known protein mixture. Two or more positive peptide filter and Unified Score filter were applied to the 22 datasets. The filters effect was evaluated according to the true positive rate of protein and peptide for each filter. RESULTS: The positive rates of protein and peptide from two or more peptide filter raised from 56.49% to 92.86%-99.12% (for protein) and from 90.67% to 97.74%-99.62% (for peptide), but many positive proteins were filtered out. The positive rates of protein and peptide from Unified Score (ThermoFinnigan value 2400) were only about 35.51% and 82.99%, but after adjusted the value (3900) according to the number of false positive peptide, those positive rate raised to 63.61% (for protein) and 91.97% (for peptide). CONCLUSIONS: Two or more peptides requirement could significantly decrease false positive rate, but it also may filter out many true positive proteins especially low molecular weight and less abundant proteins. Unified Score may be a better filter than Xcorr and DeltaCn combination and the value of 3900 is found to be more suitable for this particular datasets.

Algorithms↗

Canonical correlation analysis for data reduction in data mining applied to predictive models for breast cancer recurrence.

Data mining methods can be used for extracting specific medical knowledge such as important predictors for recurrence of breast cancer in pertinent data material. However, when there is a huge quantity of variables in the data material it is first necessary to identify and select important variables. In this study we present a preprocessing method for selecting important variables in a dataset prior to building a predictive model.In the dataset, data from 5787 female patients were analysed. To cover more predictors and obtain a better assessment of the outcomes, data were retrieved from three different registers: the regional breast cancer, tumour markers, and cause of death registers. After retrieving information about selected predictors and outcomes from the different registers, the raw data were cleaned by running different logical rules. Thereafter, domain experts selected predictors assumed to be important regarding recurrence of breast cancer. After that, Canonical Correlation Analysis (CCA) was applied as a dimension reduction technique to preserve the character of the original data.Artificial Neural Network (ANN) was applied to the resulting dataset for two different analyses with the same settings. Performance of the predictive models was confirmed by ten-fold cross validation. The results showed an increase in the accuracy of the prediction and reduction of the mean absolute error.

Breast Neoplasms↗

Robust diagnosis of non-Hodgkin lymphoma phenotypes validated on gene expression data from different laboratories.

A major challenge in cancer diagnosis from microarray data is the need for robust, accurate, classification models which are independent of the analysis techniques used and can combine data from different laboratories. We propose such a classification scheme originally developed for phenotype identification from mass spectrometry data. The method uses a robust multivariate gene selection procedure and combines the results of several machine learning tools trained on raw and pattern data to produce an accurate meta-classifier. We illustrate and validate our method by applying it to gene expression datasets: the oligonucleotide HuGeneFL microarray dataset of Shipp et al. (www.genome.wi.mit.du/MPR/lymphoma) and the Hu95Av2 Affymetrix dataset (DallaFavera's laboratory, Columbia University). Our pattern-based meta-classification technique achieves higher predictive accuracies than each of the individual classifiers , is robust against data perturbations and provides subsets of related predictive genes. Our techniques predict that combinations of some genes in the p53 pathway are highly predictive of phenotype. In particular, we find that in 80% of DLBCL cases the mRNA level of at least one of the three genes p53, PLK1 and CDK2 is elevated, while in 80% of FL cases, the mRNA level of at most one of them is elevated.

Biomarkers, Tumor↗

Fourier harmonic approach for visualizing temporal patterns of gene expression data.

DNA microarray technology provides a broad snapshot of the state of the cell by measuring the expression levels of thousands of genes simultaneously. Visualization techniques can enable the exploration and detection of patterns and relationships in a complex dataset by presenting the data in a graphical format in which the key characteristics become more apparent. The purpose of this study is to present an interactive visualization technique conveying the temporal patterns of gene expression data in a form intuitive for non-specialized end-users. The first Fourier harmonic projection (FFHP) was introduced to translate the multi-dimensional time series data into a two dimensional scatter plot. The spatial relationship of the points reflect the structure of the original dataset and relationships among clusters become two dimensional. The proposed method was tested using two published, array-derived gene expression datasets. Our results demonstrate the effectiveness of the approach.

Algorithms↗

Automated quantification of myocardial ischemia and wall motion defects by use of cardiac SPECT polar mapping and 4-dimensional surface rendering.

SPECT of cardiac perfusion and blood pools provides ungated 3-dimensional and gated 4-dimensional (4D) datasets of the ventricular myocardium. Modern reconstruction and review software is used to reorient the transverse thoracic slices into cardiac short-axis slices. Several validated algorithms are used to analyze these data. These programs segment out the left ventricle, determine the apical and basal limits, and then contour the endo- and epicardial surfaces. From these, 4D images of cardiac function that enable a dynamic review of wall motion are obtained. Global function in gated studies can be quantified by automatic computation of stroke volume and ejection fraction. Remapping of myocardial perfusion and wall motion into polar maps enables standardized quantification of the extent and severity of heart disease after comparison with databases of healthy hearts (normal databases). Several validated software packages make processing of these SPECT datasets comparatively easy and operator independent. The objectives of this review article are to describe the steps in the processing of a cardiac SPECT dataset for viewing and quantification, to explain the underlying algorithms used for automated processing, to compare the features of various software packages, to demonstrate how to read polar maps, and to identify and correct artifacts resulting from errors in automated processing.

Gated Blood-Pool Imaging↗

Consultations in general practice and at an Aboriginal community controlled health service: do they differ?

INTRODUCTION: Despite the widely acknowledged health disparities between Indigenous and non-Indigenous Australians, little is known about consultations in primary care with Indigenous people. In particular, the nature of consultations in the Aboriginal Community Controlled Health Service (ACCHS) sector has been rarely studied. Data collection about consultations in primary care has been steadily improving, with good quality data now available on an ongoing basis about patient demographics, risk factors and consultation content in private general practice. This study aimed to characterise consultations at Townsville Aboriginal and Islander Health Service (TAIHS) in terms of patient demographics and consultation content. These could then be compared with existing datasets for local consultations in mainstream general practice and from a geographically distant ACCHS. METHODS: We conducted a prospective questionnaire audit of all consultations at Townsville Aboriginal and Islander Health Service (TAIHS) over two fortnights, 6 months apart in 2000 and 2001. The questionnaire was adapted from one used in previous general practice surveys, and was completed by the treating clinician at the end of each consultation. The questionnaire described consultations using the following variables: date of consultation; patient age; ethnicity and gender; postcode and whether or not they were new to the practice; where they were seen; the provider of the service (doctor, nurse, health worker etc); Medicare level of consultation; patient reasons for encounter; problems managed; treatment and medications given; investigations; admissions; follow up; and referral. Proportions with 95% confidence intervals were calculated to facilitate comparisons with other datasets. Comparison was made with previously reported data from mainstream Townsville general practice (via the local BEACH study report) and from Darwin ACCHS (Danila Dilba). RESULTS: Of 1211 consultations studied, 1994 problems managed were recorded. TAIHS patients had a significantly younger age distribution than patients in mainstream general practice (as did patients at Danila Dilba). TAIHS consultations involved the management of more problems (1.65 problems per consultation; 95%CI [1.60, 1.70]), when compared with mainstream general practice (Townsville BEACH study 1.45 problems per consultation [1.37, 1.52]; 1.48 for Indigenous patients). Danila Dilba recorded an average of 1.58 problems managed per consultation (95% CI [1.51, 1.65]). The most frequently managed problems differed between all three datasets, and at TAIHS the most common problems managed were type 2 diabetes mellitus (11.3 times per 100 consultations), upper respiratory tract infections (9.6) and hypertension (7.9). Aboriginal Health Workers (AHW) saw the patient at TAIHS in 224/1213 (18.5%) of consultations, nurses (two Indigenous) participated in 513 (42.3%) of consultations, and a (non-Indigenous) medical officer saw the patient in 1070 (88.2%) of consultations. The Danila Dilba study found that 42.6% of their consultations involved an Aboriginal health worker only, and a health worker and a doctor managed 53.5%; only 3.9% were managed by a doctor alone without input from a health worker. CONCLUSIONS: The greater number of problems managed per consultation in ACCHS, compared with Indigenous patients in mainstream general practice, supports the assertion that ACCHS fill an important role in the health system by providing care for their largely Indigenous patients with complex care needs. The Medicare system as it was structured at the time did not encourage involvement of Indigenous health workers in provision of primary medical care. It remains to be seen whether introduction of the new enhanced primary care Medicare numbers will assist in this process. These findings have implications for ACCHS in other areas of the country and for other providers of primary health care for Indigenous Australians.

Adolescent↗

Computational strategy for discovering druggable gene networks from genome-wide RNA expression profiles.

We propose a computational strategy for discovering gene networks affected by a chemical compound. Two kinds of DNA microarray data are assumed to be used: One dataset is short time-course data that measure responses of genes following an experimental treatment. The other dataset is obtained by several hundred single gene knock-downs. These two datasets provide three kinds of information; (i) A gene network is estimated from time-course data by the dynamic Bayesian network model, (ii) Relationships between the knocked-down genes and their regulatees are estimated directly from knock-down microarrays and (iii) A gene network can be estimated by gene knock-down data alone using the Bayesian network model. We propose a method that combines these three kinds of information to provide an accurate gene network that most strongly relates to the mode-of-action of the chemical compound in cells. This information plays an essential role in pharmacogenomics. We illustrate this method with an actual example where human endothelial cell gene networks were generated from a novel time course of gene expression following treatment with the drug fenofibrate, and from 270 novel gene knock-downs. Finally, we succeeded in inferring the gene network related to PPAR-alpha, which is a known target of fenofibrate.

Bayes Theorem↗

Review of the two sample t tests.

The t test is a valuable statistical manipulation of moderate strength to determine whether a significant difference exists between the means of two groups, either paired or unpaired. Generally, the larger the t value, the greater chance of its statistical significance. Sample size also will influence the point at which t becomes significant, that is, the larger the size of n, the smaller the t required to become significant. As the number of degrees of freedom becomes larger, a smaller t value is sufficient to reject the null hypothesis, and as variability increases, the true chance for significant difference decreases. It should be noted that a common error in the use of the t test occurs with excessive repetition of the test on the same dataset. A resultant type 1 error will be introduced that incorrectly concludes that significant differences have been demonstrated when, in fact, they have resulted not from significance but from repeated statistical application. In other words, if a .05 level of significance is used repeatedly on a dataset, the investigator is assuming that there is a 1 in 20 chance of finding a significant difference when there is no true difference between variables. It becomes obvious that if 20 repeated applications of the t test were performed, one would expect to find one significant difference by chance alone when no true difference exists. If more than 10 applications of the t test are being conducted on the same dataset, the investigator should lower the level of significance (.01) on each individual t test of the entire set so fewer null hypotheses are rejected. t Tests numbering 40 to 50 should have an alpha level of .005. An alternative approach would be to consult a statistician and use multivariate statistical procedures.

Analysis of Variance↗

The registration function as a critical dependency in a lifetime clinical record (LCR).

Over the past two years, we have successfully migrated the Regenstrief Clinical Information System into our hospital. Integral to this process was the need to develop interfaces and processes supporting movements of patient identification data between the existing clinical management system (Unity, SMS) and Carebase (RCIS). Critical to the implementation of Carebase was the development of an interface between Carebase and the registration system based upon a unique medical record number. Even more critical was the development of stable processes that supported the accurate patient identification and assignment of medical record numbers. The medical record number at our institution is assigned or verified at the time of registration. Major problems occurred when patients presented during system down-times and existing medical record numbers could not be accessed, resulting in multiple registrations and medical record numbers for the same patient. This resulted in data fragmentation and required merging at a later date. Other more serious problems resulted from the assignment of the same medical record number to separate patients and with the mixing of data from multiple patients into one patient record. This was largely due to the failure of clerical personnel to appropriately identify patients at the time of registration, or multiple patients sharing identification documents, a common problem in our geographic area. Given that clinical data was to be maintained and added to the repository for several decades, errors such as these in registration would prove catastrophic. The interfaces between the various clinical systems that pass data to Carebase are all HL-standard and largely prevent data passage if registration data is inaccurate. During the early stages of implementation, approximately 300 exceptions per day were generated from clinical systems attempting to pass data to the repository. Following re¿engineering of the registration process, education of clerical personnel, and analysis of exception type, the number of exceptions due to faulty registration data fell to less than one per week. To achieve improvement in exception volume, several innovative measures were undertaken. Firstly, down-time procedures were changed to require query of the LCR for existing registration data. The LCR was maintained on a separate platform that experienced essentially no down-time and was available for this purpose. This largely eliminated the need for the use of "down-time numbers" or medical record numbers that could be temporarily assigned to patients registered when the registration system was unavailable (data would subsequently be merged into existing patient records if the patient was found to be currently in the system). If the patient was not in the LCR, then a permanent number was assigned in sequence. A registration dataset was developed and encoded onto a magnetic card (Carecard, Eltrax) and carried by patients. This enabled the rapid verification of registration data on subsequent visits to the parent institution or affiliated clinical sites. The issue of fraudulent use of the card and encoded registration dataset, however, remained problematic. Currently, a new imaging system is being installed that will soon enable the inclusion of a photograph of the patient as a component of the registration dataset. Perhaps the most significant change in the registration process involved the education of central registration and admitting personnel. An educational program was developed that reinforced the need for accuracy in collecting registration data, identifying patients, and assigning medical record numbers; more importantly, it stressed the linkage of the registration function and patient care. Lastly, an aggressive approach to monitoring exceptions resulting from errors in registration was developed. A near real-time process for identifying errors in registrations allowed for rapid intervention and feedback to involved de

Medical Records Systems, Computerized↗

The hospital information system as a source for the planning and feed-back of specialized health care.

1. INTRODUCTION. In university hospitals, choices are made to which extend specialized health care will be supported. It is characteristic, for this type of care, that it takes place in a process of the continual advance of medical technology and the growing awareness by consumers and payors. Specialized healthcare contributes to the hospital qualifiers having a political and strategic impact. The hospital board needs information for planning and budgeting these new tasks. Much of the information will be based on data stored in the Hospital Information System (HIS). Due to load limitations, instant retrieval is not preferred. A separate executive information system, uploaded with HIS data, features statistics, on a corporate level, with the power to drill-down to detailed levels. However, the ability to supply information on new types of healthcare is limited since most of these topics require a flexible system for new dedicated cross-sections, like medical treatment from several specialisms and functional levels. 2. DATA RETRIEVAL AND DISTRIBUTION. During the information analysis, details were gathered on the necessary working procedures and the administrative organization, including the data registration in the HIS. In the next phase, all relevant data was organized in a relational datamodel. For each topic of care, dedicated views were developed at both low and high aggregation levels. It revealed that a matching change of the administrative organization was required, with an emphasis on financial registration aspects. For the selection of relevant data, a bottom-up approach was applied, which was based on the registrations starting from the patient administrative subsystem, through several transactional systems, ending at the general ledger in the HIS. Data on all levels was gathered, resulting in medical details presented in quantities, up to financial figures expressed in amounts of money. This procedure distinguishes from the predefined top-down techniques generally used for management and executive information systems. Data was regularly collected from the HIS, then converted and reorganized into relational datasets using XBase protocols. After having performed central quality controls and privacy protection measures, the datasets were distributed electronically to local PCs. Standard low-cost software packages enable analyses by user-friendly selection and presentation facilities. 3. EVALUATION. The method developed for data retrieval is flexible and easy to implement. If all basic data is registered in the HIS, the procedure can be applied for all strategic hospital functions that require planning and controlling during a certain time. Critical success factors and pitfalls will be presented in the poster. Using one consistent dataset, the information required about production and budget is presented at several functional levels and is quantified in units familiar to that level. The motivation for fast and accurate registration in the HIS was improved from the moment the medical and administrative staff recognized their own data in the feed-back on specialized health care.

Decision Making, Organizational↗

An atlas of non-redundant sequences and structures of transcription factor assemblies across domains of life.

Transcription factors (TFs) regulate gene expression by controlling the recruitment of transcriptional machinery to regulatory regions of the genome. Nearly 10% of the human genome encodes TFs, making them one of the largest protein families. Despite their central roles in gene regulation, TFs are historically considered challenging therapeutic targets due to their complex interactions with DNA, RNA and associated proteins. Although recent progress in studying TFs both at molecular and structural level excels our understanding on their function, yet a universal rule decoding their recognition process remains elusive. Here, we present a curated non-redundant dataset of TFs with 3570 sequences and 377 structures. We further characterize "unique interfaces" by quantifying interface identity across interacting chains in TF assemblies. Surprisingly, our data shows that the "unique interfaces" have optimal size ranging from 2000&#x202f;&#xc5;2 to 4000&#x202f;&#xc5;2 irrespective of their quaternary assembly. To understand the functional diversity, we integrate sequence motifs, structural domains, subcellular localization and functional enrichment of TFs. We have also catalogued association of TFs with various human diseases. Our dataset provides a comprehensive platform to perform large scale analysis of TF-assemblies and aid in computational methods for their prediction across domains of life.

Gene regulation↗

Whole-genome sequencing of 490,640 UK Biobank participants.

Whole-genome sequencing provides an unbiased and complete view of the human genome and enables the discovery of genetic variation without the technical limitations of other genotyping technologies. Here we report on whole-genome sequencing of 490,640 UK Biobank participants, building on previous genotyping effort1. This advance deepens our understanding of how genetics associates with disease biology and further enhances the value of this open resource for the study of human biology and health. Coupling this dataset with rich phenotypic data, we surveyed within- and cross-ancestry genomic associations and identified novel genetic and clinical insights. Although most associations with disease traits were primarily observed in individuals of European ancestries, strong or novel signals were also identified in individuals of African and Asian ancestries. With the improved ability to accurately genotype structural variants and exonic variation in both coding and UTR sequences, we strengthened and revealed novel insights relative to whole-exome sequencing2,3 analyses. This dataset, representing a large collection of whole-genome sequencing&#xa0;data that is available to the UK Biobank research community, will enable advances of our understanding of the human genome, facilitate the discovery of diagnostics and&#xa0;therapeutics with higher efficacy and improved safety profile, and enable precision medicine strategies with the potential to improve global health.

Humans↗

Microbial genomic database of the Yangtze River, the third-longest river on Earth.

Microbes play an important role in mediating the nutrient cycling in the river ecosystem as a hotspot for biogeochemical processes. Due to scattered sampling efforts, however, there is a lack of a systematic study of the diversity of prokaryotic genomes in the Yangtze River, the third longest river on Earth. Here, we collected 602 metagenomic datasets of water, sediment and riparian soil samples spanning the Upper, Middle, and Lower basins of the Yangtze River over a 6,300&#x2009;km continuum. We reconstructed 8,110 qualified genomes represented by 927 species-level genomes at the 95% ANI threshold, spanning 31 bacterial and five archaeal phyla. We further showed that more than half of these species (61.3% ~ 82.4%) were novel according to the genomic comparison against the curated databases, greatly expanding the known diversity of river prokaryotes. This dataset depicts an overview of microbial genomic diversity in the Yangtze River and provides a resource for in-depth investigation of metabolic potential, ecology, and evolution of riverine microbiomes.

Rivers↗

Microbial metagenomes from Lake Soyang, the largest freshwater reservoir in South Korea.

Lake ecosystems play a fundamental role in the global biogeochemical cycling of essential elements such as carbon, nitrogen, and phosphorus. Microorganisms within these ecosystems mediate key processes that regulate these cycles. Metagenomic analyses provide valuable insights into the taxonomic and functional diversity of microbial communities in various environments, including freshwater habitats. Here, we present a comprehensive metagenomic dataset derived from Lake Soyang, the largest freshwater reservoir in South Korea. A total of 28 metagenomes were generated from water samples collected across two distinct sampling periods: the first set (n&#x2009;=&#x2009;8) was obtained between April 2014 and January 2015 from two depths (1&#x2009;m and 50&#x2009;m) in four different seasons, while the second set (n&#x2009;=&#x2009;20) was collected between January 2019 and November 2019 from five depths (1, 10, 20, 40, and 90&#x2009;m) over four seasons. Metagenomic sequencing yielded 9.3-21.8 Gbp per sample. This dataset provides a valuable resource for future studies exploring the ecophysiological characteristics of microbial communities in pelagic freshwater environments.

Republic of Korea↗

A compendium of horizontal gene transfers in Metazoa.

With more eukaryotic genomes available for study researchers have been able to identify a growing number of horizontal gene transfer (HGT) candidates. We compiled 9,495 protein coding genes that were identified as horizontally transferred to metazoan hosts in the published literature. This dataset contains gene transfers from bacteria, fungi, archaea and protists to metazoans. We assigned a confidence score to each gene based on the methods used in the scientific paper reporting HGT. All the coding sequences and protein sequences for the HGT genes are stored in a fig share repository. This dataset can be used to identify trends in genome and protein evolution and provide a foundation for creating a centralized HGT database for eukaryotes.

Gene Transfer, Horizontal↗