Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Datasets as Topic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

The limits of competing interest disclosures.

OBJECTIVE: To assess the effectiveness of conflict of interest disclosure policies by comparing a competing interests disclosure statement that met the requirements established by the journal in a 2003 article on health effects of secondhand smoke based on the American Cancer Society CPS-I dataset with internal tobacco industry documents describing financial ties between the tobacco industry and authors of the study. DESIGN: Descriptive analysis of internal tobacco industry documents retrieved from the Legacy Tobacco Documents Library, University of California, San Francisco. RESULTS: Meeting the requirements for financial disclosure established by the journal did not provide the reader with a full picture of the tobacco industry's involvement with the study authors. The tobacco industry documents reveal that the authors had long standing financial and other working relationships with the tobacco industry. CONCLUSION: These findings are another example of how simply requiring authors to disclose financial ties with the tobacco industry may not be adequate to give readers (and reviewers) a full picture of the author's relationship with the tobacco industry. The documents also reveal that the industry funds research to enhance its credibility and endeavours to work with respected scientists to advance its goals. These findings question the adequacy of current journal policies regarding competing interest disclosures and the acceptability of tobacco industry funding for academic research.

Biomedical Research↗

Meta-analysis models with group structure for pleiotropy detection at gene and variant level using summary statistics from multiple datasets.

Genome-wide association studies (GWASs) have highlighted the importance of pleiotropy in human diseases, where one gene can impact 2 or more unrelated traits. Examining shared genetic risk factors across multiple diseases can enhance our understanding of these conditions by pinpointing new genes and biological pathways involved. Furthermore, with an increasing wealth of GWAS summary statistics available to the scientific community, leveraging these findings across multiple phenotypes could unveil novel pleiotropic associations. Existing selection methods examine pleiotropic associations one by one at a scale of either the genetic variant or the gene, and thus cannot consider all the genetic information at the same time. To address this limitation, we propose a new approach called MPSG (Meta-analysis model adapted for Pleiotropy Selection with Group structure). This method performs a penalized multivariate meta-analysis method adapted for pleiotropy and takes into account the group structure information nested in the data to select relevant variants and genes (or pathways) from all the genetic information. To do so, we implemented an alternating direction method of multipliers algorithm. We compared the performance of the method with other benchmark meta-analysis approaches such as GCPBayes, PLACO, and ASSET by considering as inputs different kinds of summary statistics. We provide an application of our method to the identification of potential pleiotropic genes between breast and thyroid cancers.

Humans↗

Sensitivity and specificity of a screening questionnaire for dry eye.

We developed a Dry Eye Screening Questionnaire for the Dry Eye Epidemiology Projects (DEEP), a proposed large epidemiologic study. All persons who screen positive and a small sample of those who screen negative are to be invited for a diagnostic examination. Containing 19 questions, of which only 14 were used in the analysis, the questionnaire takes only a few minutes to administer on the telephone. To construct a discriminator function and thus a ROC curve, we used stepwise multiple regression on screening responses from a clinic series of 77 cases and 79 controls. Stepwise regression may incorporate into the predictor equation variables whose relation to the predicted is only accidental. Further, misclassification rates are underestimated by the resubstitution method, in which the proportion misclassified is obtained from the same dataset in which the discriminator function was fitted. To counter these problems, we randomly divided the data in half. We chose as predictors only those variables (Dry and Irritated) selected by stepwise regression in both data halves. We estimated unbiased misclassification rates using the unbiased test set method, in which the discriminator is fitted in one data half, and misclassification rates are calculated in the other half. Comparison of ROC curves arising from resubstitution and test set estimates indicates that resubstitution bias in misclassification rate estimation is negligible in our data. A resubstitution estimate made on the entire data is thus preferred. The resulting sensitivity/specificity values are reasonably high (e.g., 60%/94%), suggesting that the questionnaire will be a useful screening tool in the DEEP study. A second discriminator using the sum of all 14 responses is similar in its misclassification characteristics to the first discriminator. A second potentially significant error, arising from applying results from a clinical series to a general population, will be investigated as survey results in DEEP become available.

Adult↗

Visual analysis of serial T2-weighted MRI in multiple sclerosis: intra- and interobserver reproducibility.

We evaluated the effect of consensus formation and training on the agreement between observers in scoring the number of new and enlarging multiple sclerosis (MS) lesions on serial T2-weighted MRI studies. The baseline and month 9 MRI studies of 16 patients with a range of MRI activity were used (dual-echo conventional spin-echo sequence, TR 2000, TE 34 and 90 ms, 5 mm contiguous slices, inplane resolution 1 mm). First, the serial studies were visually analysed for the presence of new and enlarging lesions, on two occasions, by five experienced observers, without adopting any consensus strategy and in isolation. Next, the observers met to identify the common sources of inconsistencies in reporting between observers and formulate consensus rules. Finally, a further independent reading session was performed on the same MRI dataset, this time applying the consensus rules. Agreement between observers was assessed using kappa scores. Without the consensus rules, interobserver kappa scores for the first and second reading sessions for new lesions were only 0.51 and 0.39 respectively; agreement for enlarging lesions was even worse. The mean intraobserver kappa score for new lesions was higher at 0.72, reflecting the fact that the observers were consistently applying their individual assessment strategies. Application of the consensus rules did not lead to a significant improvement in inter observer kappas; the kappa scores adopting the guidelines were 0.46 and 0.21 for new and enlarging lesions respectively. Consensus guidelines thus did not improve the reproducibility of visual analysis of serial T2-weighted MRI, and the level of agreement between observers remained only moderate. Suboptimal repositioning is likely to be a major source of residual variability and this suggests a future role for image registration strategies; until then, a single observer, or pair of observers working in consensus, should be used in MS studies.

Analysis of Variance↗

Use of clinical syndromes to target antibiotic prescribing in seriously ill children in malaria endemic area: observational study.

OBJECTIVES: To determine how well antibiotic treatment is targeted by simple clinical syndromes and to what extent drug resistance threatens affordable antibiotics. DESIGN: Observational study involving a priori definition of a hierarchy of syndromic indications for antibiotic therapy derived from World Health Organization integrated management of childhood illness and inpatient guidelines and application of these rules to a prospectively collected dataset. SETTING: Kilifi District Hospital, Kenya. PARTICIPANTS: 11,847 acute paediatric admissions. MAIN OUTCOME MEASURES: Presence of invasive bacterial infection (bacteraemia or meningitis) or Plasmodium falciparum parasitaemia; antimicrobial sensitivities of isolated bacteria. RESULTS: 6254 (53%) admissions met criteria for syndromes requiring antibiotics (sick young infants; meningitis/encephalopathy; severe malnutrition; very severe, severe, or mild pneumonia; skin or soft tissue infection): 672 (11%) had an invasive bacterial infection (80% of all invasive bacterial infections identified), and 753 (12%) died (93% of all inpatient deaths). Among P falciparum infected children with a syndromic indication for parenteral antibiotics, an invasive bacterial infection was detected in 4.0-8.8%. For the syndrome of meningitis/encephalopathy, 96/123 (76%) isolates were fully sensitive in vitro to penicillin or chloramphenicol. CONCLUSIONS: Simple clinical syndromes effectively target children admitted with invasive bacterial infection and those at risk of death. Malaria parasitaemia does not justify withholding empirical parenteral antibiotics. Lumbar puncture is critical to the rational use of antibiotics.

Anti-Bacterial Agents↗

Deciphering Arabidopsis thaliana gene neighborhoods through bibliographic co-citations.

In the framework of genome annotation, scientific literature is obviously the major source of biological knowledge. The aim of the work described in this paper is to exploit this source of data for the model plant Arabidopsis thaliana. The first step has consisted in constituting a relevant bibliographic references dataset for plant genomic research. Genes co-citations have then been systematically annotated in this reference dataset, starting from the simple idea that if genes are cited in the same publication, they must probably share some related functional properties. In order to deal with the synonymous gene name problem, a gene name reference list has been constituted starting from A. thaliana SwissProt entries. This list was used to build clusters of co-cited genes by a single linkage procedure such that any gene in a given cluster possesses at least one co-cited partner in the same cluster. Analysis of the clusters demonstrate the biological consistency of this approach, with only very few fortuitous links. As an example, a cluster including genes related to flowering time is more deeply described in the paper. Finally, a graphical representation of each cluster was performed, which provides a convenient way to retrieve the genes (the nodes of the graphs) and the references in which they were co-cited (the edges of the graphs). All the results can be accessed at the URL http://chlora.Igi.infobiogen.fr:1234/bib_arath/.

Arabidopsis↗

The validity of the Neale and Kendler model-fitting approach in examining the etiology of comorbidity.

Given that knowledge regarding the etiology of comorbidity between disorders can have a significant impact on research regarding the classification, treatment, and etiology of the disorders, the ability to reject incorrect hypotheses regarding the causes of comorbidity is very important. A simulation study was conducted to assess the validity of the Neale and Kendler (1995) model-fitting approach in examining the etiology of comorbidity between two disorders. First, data were simulated under the assumptions of the 13 alternative comorbidity models described by Neale and Kendler. Second, model-fitting analyses testing the comorbidity models were conducted on the simulated datasets. Thirteen sets of data with varying model parameters were simulated to test Neale and Kendler's assertion that their model-fitting approach is appropriate across a range of potential prevalences and degrees of familiality. The validity of the model-fitting approach in examining unselected twin data and a combination of selected family data and unselected family data was explored. The model-fitting approach successfully discriminated several classes of comorbidity models, although discrimination between models within classes of related models was less accurate. Results suggest that the model-fitting approach can be a useful tool in examining the etiology of the comorbidity between disorders if the caveats of the present study's results are considered carefully. As predicted by Neale and Kendler, variations in the disorder prevalences and familial correlations did not affect the validity of their model-fitting approach, but affected the power to discriminate the correct model. As suggested by Neale and Kendler, the model-fitting approach can be applied to both unselected and selected data and to both twin and family data.

Comorbidity↗

Relationship among hospital ERCP volume, length of stay, and technical outcomes.

BACKGROUND: The relationship between hospital procedure volume and outcome has been recognized for various specialties and procedures. Although increasingly used and in existence for 40 years, to date, data on the relationship between hospital volume and outcome of ERCP are scant. OBJECTIVE: We sought to examine health-related outcomes after ERCP in relation to hospital procedure volume. DESIGN: Secondary analysis of a national administrative database. We used the National Inpatient Sample (NIS) database to evaluate health-related outcomes among patients who underwent ERCP from 1998 to 2001. MAIN OUTCOME MEASUREMENTS: Logistic and multiple regression models were used to estimate the association of hospital ERCP volume with length of stay (LOS), rates of procedural failure, and mortality. Fixed effect models were used to adjust for all time invariant hospital characteristics for each hospital within the dataset. RESULTS: Data from 2629 hospitals that performed 199,625 ERCPs were evaluated. The median number of ERCPs performed in participating hospitals was 49 per year (range, 1-1004), with 25% of hospitals performing > or =100 ERCPs per year and 5% performing > or =200 per year. Significant trends in the relationship between volume and outcome were observed with respect to LOS and procedural failure: the median LOS was lower in high-volume (> or =200 ERCP/y) than low-volume (< or =100 ERCP/y) hospitals (6.9 vs 7.8 days, p < 0.0001) and the mean difference in expected LOS was 1.08 days (p < 0.0001). Multivariate regressions with hospital level fixed effects found significant negative relationships between procedure volume and procedure failure rates, but no significant effect on inpatient mortality rates was detected. LIMITATIONS: NIS database permits analyses of only inpatient ERCPs. It precludes analysis of procedural complications, reinterventions, and influence of individual provider volume on outcomes. CONCLUSIONS: Inpatients who undergo ERCP at high-volume hospitals have shorter LOS and lower procedural failure rates than those undergoing ERCP at low-volume hospitals. These findings have important implications for health care policy decision making and resource utilization.

Cholangiopancreatography, Endoscopic Retrograde↗

Deriving sediment quality guidelines from field-based species sensitivity distributions.

The determination of predicted no-effect concentrations (PNECs) and sediment quality guidelines (SQGs) of toxic chemicals in marine sediment is extremely important in ecological risk assessment. However, current methods of deriving sediment PNECs or threshold effect levels (TELs) are primarily based on laboratory ecotoxicity bioassays that may not be ecologically and environmentally relevant. This study explores the possibility of utilizing field data of benthic communities and contaminant loadings concurrently measured in sediment samples collected from the Norwegian continental shelf to derive SQGs. This unique dataset contains abundance data for ca. 2200 benthic species measured at over 4200 sampling stations, along with co-occurring concentration data for >25 chemical species. Using barium, cadmium, and total polycyclic aromatic hydrocarbons (PAHs) as examples, this paper describes a novel approach that makes use of the above data set for constructing field-based species sensitivity distributions (f-SSDs). Field-based SQGs are then derived based on the f-SSDs and HCx values [hazardous concentration for x% of species or the (100-x)% protection level] by the nonparametric bootstrap method. Our results for Cd and total PAHs indicate that there are some discrepancies between the SQGs currently in use in various countries and our field-data-derived SQGs. The field-data-derived criteria appear to be more environmentally relevant and realistic. Here, we suggest that the f-SSDs can be directly used as benchmarks for probabilistic risk assessment, while the field-data-derived SQGs can be used as site-specific guidelines or integrated into current SQGs.

Animals↗

Cesarean births in Taiwan.

OBJECTIVE: To evaluate the use of cesarean delivery in Taiwan by comparing local clinical indications with those in international cohorts. METHODS: In-patient claims from the National Health Insurance (NHI) in Taiwan were analyzed. Indications for cesarean delivery were evaluated with primary diagnosis codes and procedure codes from the NHI dataset. To produce a stable numerator for cesarean section, 3 years (1998-2000) of claims for cesarean delivery were abstracted and annualized. RESULTS: Rates ranged between 27.3% and 28.7% for primary cesarean delivery and were below 5% for vaginal birth after a cesarean section (VBAC). Compared with rates in other countries, rates for overall and primary cesarean section as well as for VBAC were significantly higher in medical centers in Taiwan (P<0.001). However, the clinics contributed the most to the difference in both overall and primary cesarean rates. The most common indication for cesarean section was prior cesarean section (43.3%-45.5%), followed by malpresentation (19.6%-23.4%). The proportion of fetuses with malpresentation delivered by cesarean section in Taiwan was 7.9%, almost twice the upper limit expected for all pregnancies as indicated in international studies. CONCLUSION: It is important to use appropriately documented data and to compare them with international data when monitoring local obstetric practices. The disproportionately high cesarean delivery rates in Taiwan may hold major lessons for the many countries contemplating or having universal health insurance coverage with a similar mix of providers.

Birthing Centers↗

How many work-related injuries requiring hospitalization in British Columbia are claimed for workers' compensation?

BACKGROUND: Workplace compensation claims datasets represent an important source of information on work-related injuries. This study investigated the concordance between hospital discharge records and workers' compensation records for work-related serious injuries among a cohort of sawmill workers in British Columbia (BC), Canada. It also examined the extent to which workers' compensation capturing patterns varied by cause, severity of injuries, and demographic characteristics of workers. METHODS: Work-related injuries were identified in hospitalization records between April 1989 and December 1998, and were matched by dates and description of injury to compensation records. RESULTS: The agreement between the hospital records and compensation records was good (kappa = 0.84, P < 0.01). A lower claim reporting rate for work-related hospitalization was observed for older and non-white workers. More serious injuries defined by longer length of stay and emergency admissions were more likely to be reported. Falls, struck against, and overexertion injuries had lower reporting rates; whereas, machinery-related, cutting/piercing, and caught in/between injuries had higher reporting rates. CONCLUSIONS: When compared with hospital discharge records, the compensation agency underreported incidents of serious work-related injuries by 10-15% among the sawmill workers.

Adult↗

Evaluation of a transit first-aid station providing emergency care to former Yugoslavian war victims evacuated in Ancona, Italy.

BACKGROUND: A first-aid station was implemented in Falconara Marittima airport (Ancona, Italy). It provided medical emergency care to war victims evacuated from former Yugoslavia in transit for further treatment. MATERIALS AND METHODS: A descriptive analysis of the displaced population arriving at the first-aid station was performed using three independent datasets for administrative information, of which one included medical information. The implemented resources were also evaluated. RESULTS: From August 1993 to March 1995, 2272 displaced persons were registered at the first-aid station, out of which 54.2% were accompanying family members. Among those needing medical intervention (45.8% of total), most frequent diagnoses were traumatisms and burns (59.8%), neoplasms (15.6%), and congenital malformations (13.2%). The medical care provided at the first-aid station was most often basic: a medical examination alone was performed on 77.0% of the patients, and a minor dressing on 17.3%. Median length of stay was 1 day. Patients were sent to 30 different countries and 8% were forwarded to the local regional hospital. Deployed logistical resources exceeded by far actual needs but a lack of psychological assistance was observed, mainly for children. The agencies involved did not coordinate data sharing and follow-up information. CONCLUSIONS: The medical assistance to the war victims was efficient regarding provided care and timeliness. Effectiveness of such a programme could be improved by a better coordination between partners, allowing more adequate logistics according to appropriate epidemiological information.

Adolescent↗

Approaches for linking whole-body fish tissue residues of mercury or DDT to biological effects thresholds.

A variety of methods have been used by numerous investigators attempting to link tissue concentrations with observed adverse biological effects. This paper is the first to evaluate in a systematic way different approaches for deriving protective (i.e., unlikely to have adverse effects) tissue residue-effect concentrations in fish using the same datasets. Guidelines for screening papers and a set of decision rules were formulated to provide guidance on selecting studies and obtaining data in a consistent manner. Paired no-effect (NER) and low-effect (LER) whole-body residue concentrations in fish were identified for mercury and DDT from the published literature. Four analytical approaches of increasing complexity were evaluated for deriving protective tissue residues. The four methods were: Simple ranking, empirical percentile, tissue threshold-effect level (t-TEL), and cumulative distribution function (CDF). The CDF approach did not yield reasonable tissue residue thresholds based on comparisons to synoptic control concentrations. Of the four methods evaluated, the t-TEL approach best represented the underlying data. A whole-body mercury t-TEL of 0.2 mg/kg wet weight, based largely on sublethal endpoints (growth, reproduction, development, behavior), was calculated to be protective of juvenile and adult fish. For DDT, protective whole-body concentrations of 0.6 mg/kg wet weight in juvenile and adult fish, and 0.7 mg/kg wet weight for early life-stage fish were calculated. However, these DDT concentrations are considered provisional for reasons discussed in this paper (e.g., paucity of sublethal studies).

Animals↗

Effects of information and machine learning algorithms on word sense disambiguation with small datasets.

Current approaches to word sense disambiguation use (and often combine) various machine learning techniques. Most refer to characteristics of the ambiguity and its surrounding words and are based on thousands of examples. Unfortunately, developing large training sets is burdensome, and in response to this challenge, we investigate the use of symbolic knowledge for small datasets. A naïve Bayes classifier was trained for 15 words with 100 examples for each. Unified Medical Language System (UMLS) semantic types assigned to concepts found in the sentence and relationships between these semantic types form the knowledge base. The most frequent sense of a word served as the baseline. The effect of increasingly accurate symbolic knowledge was evaluated in nine experimental conditions. Performance was measured by accuracy based on 10-fold cross-validation. The best condition used only the semantic types of the words in the sentence. Accuracy was then on average 10% higher than the baseline; however, it varied from 8% deterioration to 29% improvement. To investigate this large variance, we performed several follow-up evaluations, testing additional algorithms (decision tree and neural network), and gold standards (per expert), but the results did not significantly differ. However, we noted a trend that the best disambiguation was found for words that were the least troublesome to the human evaluators. We conclude that neither algorithm nor individual human behavior cause these large differences, but that the structure of the UMLS Metathesaurus (used to represent senses of ambiguous words) contributes to inaccuracies in the gold standard, leading to varied performance of word sense disambiguation techniques.

Algorithms↗

The effectiveness of human impact assessment in the Finnish Healthy Cities Network.

OBJECTIVES: To develop a framework for analysing the effectiveness of prospective assessment and to apply the framework to human impact assessments (HuIA) carried out in the Finnish Healthy Cities Network. METHODS: The framework was formed by synthesizing and developing the themes that emerged from the published literature on effectiveness. The research material consists of interviews with people who participated in the assessment process in the municipalities (19 interviews). The research material also included assessment documents, proceedings of working meetings, municipal policy documents, background material and project reports produced in the municipalities studied. The research datasets were examined by content analysis. RESULTS: HuIA increased the decision-makers' awareness of effects and functioned as a tool for empowerment. The latter was apparent, for instance, in the social welfare and healthcare sector, finding a role for itself in decisively co-ordinating interdisciplinary work and actively seeking to alleviate identified negative effects. The assessment process also opened up the planning process, committed various actors to the decision, helped select the right alternative and promoted social learning. CONCLUSIONS: From the viewpoint of preparation and decision-making, the effectiveness of a HuIA increases when assessment becomes a recurring process and an integral part of an organization's activities. Integration of an assessment into permanent structures or activities, such as drawing up programmes or preparing strategies, helps the results of the assessment to be seen more clearly. From the viewpoint of decision-making, it is also important to strengthen the decision-makers' expertise in prospective assessment. When the effectiveness of HuIA is looked at in a new way (i.e. from the viewpoint of goal achievement, decision-making or learning), a more comprehensive interpretation can be given.

City Planning↗

A detailed examination of the clinical terms and concepts required for communication by electronic messages in diabetes care.

The wider electronic exchange of clinical information between heterogeneous information systems in the delivery of diabetes care demands a common structure in the form of a message standard. A European Standard electronic diabetes message is being developed in conjunction with CEN TC251. This paper describes the methodologies that the 1998 DO IT Workshop has used to identify potential areas of difficulty in the design and implementation of the preliminary message model. To facilitate implementation and to avoid ambiguity in electronic messaging it is particularly important that there is standardisation of the definitions of the clinical terms specifically used in diabetes care across systems. Comprehensive lists of such terms to describe all areas of diabetes care do not exist and there is a lack of harmonisation of definitions in many areas. Thus, to better understand the user requirements of diabetes messaging several approaches were adopted. A review of the clinical terms and concepts contained in pre-existing datasets was undertaken with detailed study of a number of specific areas of diabetes care, analysing the conceptual structure of all the clinical terms that they comprised. Consideration of several worst case clinical scenarios for messages to communicate was also made to identify deficiencies in the message structure. This activity confirmed the importance of creating a Standard for a superset or thesaurus of diabetes specific terms, with appropriate definitions, to harmonise data communication in different IT systems to facilitate messaging. A substantial number of new terms were identified in the workshop and these will form an important first step to accomplishing a first draft superset once fully analysed. It was also apparent that certain specific areas within diabetes care, but most particularly in nursing, dietetics and podiatry, need urgent work to further develop the concepts and terms. This needs to be facilitated for an appropriate group of such professionals. To achieve such a Standard, continued co-operation with CEN/ISSS was recognised to be very important.

Communications Media↗

Purdue ionomics information management system. An integrated functional genomics platform.

The advent of high-throughput phenotyping technologies has created a deluge of information that is difficult to deal with without the appropriate data management tools. These data management tools should integrate defined workflow controls for genomic-scale data acquisition and validation, data storage and retrieval, and data analysis, indexed around the genomic information of the organism of interest. To maximize the impact of these large datasets, it is critical that they are rapidly disseminated to the broader research community, allowing open access for data mining and discovery. We describe here a system that incorporates such functionalities developed around the Purdue University high-throughput ionomics phenotyping platform. The Purdue Ionomics Information Management System (PiiMS) provides integrated workflow control, data storage, and analysis to facilitate high-throughput data acquisition, along with integrated tools for data search, retrieval, and visualization for hypothesis development. PiiMS is deployed as a World Wide Web-enabled system, allowing for integration of distributed workflow processes and open access to raw data for analysis by numerous laboratories. PiiMS currently contains data on shoot concentrations of P, Ca, K, Mg, Cu, Fe, Zn, Mn, Co, Ni, B, Se, Mo, Na, As, and Cd in over 60,000 shoot tissue samples of Arabidopsis (Arabidopsis thaliana), including ethyl methanesulfonate, fast-neutron and defined T-DNA mutants, and natural accession and populations of recombinant inbred lines from over 800 separate experiments, representing over 1,000,000 fully quantitative elemental concentrations. PiiMS is accessible at www.purdue.edu/dp/ionomics.

Arabidopsis↗

The olfactory receptor gene superfamily: data mining, classification, and nomenclature.

The vertebrate olfactory receptor (OR) subgenome harbors the largest known gene family, which has been expanded by the need to provide recognition capacity for millions of potential odorants. We implemented an automated procedure to identify all OR coding regions from published sequences. This led us to the identification of 831 OR coding regions (including pseudogenes) from 24 vertebrate species. The resulting dataset was subjected to neighbor-joining phylogenetic analysis and classified into 32 distinct families, 14 of which include only genes from tetrapodan species (Class II ORs). We also report here the first identification of OR sequences from a marsupial (koala) and a monotreme (platypus). Analysis of these OR sequences suggests that the ancestral mammal had a small OR repertoire, which expanded independently in all three mammalian subclasses. Classification of "fish-like" (Class I) ORs indicates that some of these ancient ORs were maintained and even expanded in mammals. A nomenclature system for the OR gene superfamily is proposed, based on a divergence evolutionary model. The nomenclature consists of the root symbol 'OR', followed by a family numeral, subfamily letter(s), and a numeral representing the individual gene within the subfamily. For example, OR3A1 is an OR gene of family 3, subfamily A, and OR7E12P is an OR pseudogene of family 7, subfamily E. The symbol is to be preceded by a species indicator. We have assigned the proposed nomenclature symbols for all 330 human OR genes in the database. A WWW tool for automated name assignment is provided.

Animals↗