Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Curation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Rat Genome Database (RGD): mapping disease onto the genome.

The Rat Genome Database (RGD, http://rgd.mcw.edu) is an NIH-funded project whose stated mission is 'to collect, consolidate and integrate data generated from ongoing rat genetic and genomic research efforts and make these data widely available to the scientific community'. In a collaboration between the Bioinformatics Research Center at the Medical College of Wisconsin, the Jackson Laboratory and the National Center for Biotechnology Information, RGD has been created to meet these stated aims. The rat is uniquely suited to its role as a model of human disease and the primary focus of RGD is to aid researchers in their study of the rat and in applying their results to studies in a wider context. In support of this we have integrated a large amount of rat genetic and genomic resources in RGD and these are constantly being expanded through ongoing literature and bulk dataset curation. RGD version 2.0, released in June 2001, includes curated data on rat genes, quantitative trait loci (QTL), microsatellite markers and rat strains used in genetic and genomic research. VCMap, a dynamic sequence-based homology tool was introduced, and allows researchers of rat, mouse and human to view mapped genes and sequences and their locations in the other two organisms, an essential tool for comparative genomics. In addition, RGD provides tools for gene prediction, radiation hybrid mapping, polymorphic marker selection and more. Future developments will include the introduction of disease-based curation expanding the curated information to cover popular disease systems studied in the rat. This will be integrated with the emerging rat genomic sequence and annotation pipelines to provide a high-quality disease-centric resource, applicable to human and mouse via comparative tools such as VCMap. RGD has a defined community outreach focus with a Visiting Scientist program and the Rat Community Forum, a web-based forum for rat researchers and others interested in using the rat as an experimental model. Thus, RGD is not only a valuable resource for those working with the rat but also for researchers in other model organisms wishing to harness the existing genetic and physiological data available in the rat to complement their own work.

Animals↗

Unveiling non-small cell lung cancer treatment effect heterogeneity: a comparative analysis of statistical methods.

BACKGROUND: For patients with advanced non-small cell lung cancer lacking targetable genomic alterations, the impact of clinicogenomic characteristics on the effectiveness of combining chemotherapy with immunotherapy is unclear. METHODS: We evaluated 4 statistical methods for detecting heterogeneous treatment effects related to clinical factors, including programmed death-ligand 1 expression, tumor mutation burden, and stage at diagnosis, using the American Association for Cancer Research Project Genomics Evidence Neoplasia Exchange BioPharma Collaborative dataset supplemented with institutional data collected under the same data curation model. A 2-sided P value of no more than .05 was used to denote statistical significance for all analyses. RESULTS: The mixture model revealed 2 latent subgroups: in one subgroup, there was no meaningful treatment effect, with average progression-free survival (PFS) only 5% longer with immunotherapy alone (95% confidence interval [CI] = -19% to 35%); in the second subgroup, immunotherapy alone was associated with a 35% decrease in average PFS (95% CI = -59% to 2%), corresponding to a ratio in treatment effects of 1.62 (95% CI = 1.02 to 2.57). There was a marginal association between lower tumor mutation burden levels and membership in the subgroup with improved PFS following receipt of chemoimmunotherapy. The causal survival forest highlighted the importance of tumor mutation burden (variable importance ranking: 1) and programmed death-ligand 1 (variable importance ranking: 3) when assessing heterogeneity. In contrast, the accelerated failure time and Cox proportional hazards models did not detect any statistically significant heterogeneous treatment effects. In simulations, the mixture model identified heterogeneous treatment effects more frequently than other methods, especially with weak covariate relationships, demonstrating its utility for informing personalized treatment approaches. CONCLUSIONS: The application of novel statistical methods to large scale clinico-genomic databases offers an opportunity to more accurately identify heterogeneous treatment effects in some settings as compared to traditional statistical methods. Applying such methods to the AACR Project GENIE BPC non-small cell lung cancer data indicated a potential association between decreasing tumor mutation burden and improved outcomes with chemoimmunotherapy as compared to immunotherapy alone.

Humans↗

Prediction of coordination number and relative solvent accessibility in proteins.

Knowing the coordination number and relative solvent accessibility of all the residues in a protein is crucial for deriving constraints useful in modeling protein folding and protein structure and in scoring remote homology searches. We develop ensembles of bidirectional recurrent neural network architectures to improve the state of the art in both contact and accessibility prediction, leveraging a large corpus of curated data together with evolutionary information. The ensembles are used to discriminate between two different states of residue contacts or relative solvent accessibility, higher or lower than a threshold determined by the average value of the residue distribution or the accessibility cutoff. For coordination numbers, the ensemble achieves performances ranging within 70.6-73.9% depending on the radius adopted to discriminate contacts (6A-12A). These performances represent gains of 16-20% over the baseline statistical predictor, always assigning an amino acid to the largest class, and are 4-7% better than any previous method. A combination of different radius predictors further improves performance. For accessibility thresholds in the relevant 15-30% range, the ensemble consistently achieves a performance above 77%, which is 10-16% above the baseline prediction and better than other existing predictors, by up to several percentage points. For both problems, we quantify the improvement due to evolutionary information in the form of PSI-BLAST-generated profiles over BLAST profiles. The prediction programs are implemented in the form of two web servers, CONpro and ACCpro, available at http://promoter.ics.uci.edu/BRNN-PRED/.

Amino Acids↗

Large-scale prediction of disulphide bridges using kernel methods, two-dimensional recursive neural networks, and weighted graph matching.

The formation of disulphide bridges between cysteines plays an important role in protein folding, structure, function, and evolution. Here, we develop new methods for predicting disulphide bridges in proteins. We first build a large curated data set of proteins containing disulphide bridges to extract relevant statistics. We then use kernel methods to predict whether a given protein chain contains intrachain disulphide bridges or not, and recursive neural networks to predict the bonding probabilities of each pair of cysteines in the chain. These probabilities in turn lead to an accurate estimation of the total number of disulphide bridges and to a weighted graph matching problem that can be addressed efficiently to infer the global disulphide bridge connectivity pattern. This approach can be applied both in situations where the bonded state of each cysteine is known, or in ab initio mode where the state is unknown. Furthermore, it can easily cope with chains containing an arbitrary number of disulphide bridges, overcoming one of the major limitations of previous approaches. It can classify individual cysteine residues as bonded or nonbonded with 87% specificity and 89% sensitivity. The estimate for the total number of bridges in each chain is correct 71% of the times, and within one from the true value over 94% of the times. The prediction of the overall disulphide connectivity pattern is exact in about 51% of the chains. In addition to using profiles in the input to leverage evolutionary information, including true (but not predicted) secondary structure and solvent accessibility information yields small but noticeable improvements. Finally, once the system is trained, predictions can be computed rapidly on a proteomic or protein-engineering scale. The disulphide bridge prediction server (DIpro), software, and datasets are available through www.igb.uci.edu/servers/psss.html.

Amino Acid Sequence↗

[Status of gastrectomy in the multi-modality therapy concept of primary non-Hodgkin's lymphoma of the stomach].

Retrospectively analyzed data of 41 patients with gastric non-Hodgkin lymphoma are presented with special regard to the required extent of gastric resection in multimodality treatment. Thirty lymphomas of low, 2 of intermediate and 9 of high grade malignancy were distributed on stage Ie in 44%, stage IIe in 15%, stage IIIe in 12% and stage IV in 29%. The cumulative 5 years survival rate was 85% for stage Ie and 55% for the stages IIe and IIIe. Stage IV showed a 3 years survival of 10%. The proximal part of the stomach was involved in 73%, polycentric lesions were found in 15%. The majority (71%) of the tumours showed an infiltrating growth, an invasion of the oesophagus and/or the duodenum was seen in 20%. Five patients (12%) had synchronous adenocarcinoma as second gastric tumour. Regarding our morphological and topographical data curative surgery for primary gastric lymphoma required total gastric resection.

Adult↗

Mitogenomic and phylogenomic analyses identify a cohesive Western Atlantic lineage within the Narcine complex (Torpediniformes: Narcinidae).

BACKGROUND: Accurate species delimitation within electric rays of the genus Narcine has been hindered by overlapping morphological characters and limited molecular resolution in previous single-locus studies. This study aims to evaluate phylogenetic relationships and species boundaries within the Narcine species complex across the Western Atlantic using complete mitochondrial genomes. METHODS AND RESULTS: Seven complete mitogenomes were newly assembled from individuals representing distinct morphotypes sampled across geographically widespread Western Atlantic localities and analyzed together with publicly available reference sequences. Mitochondrial protein-coding genes (PCGs) were examined using concatenated nucleotide and amino acid datasets under partitioned maximum-likelihood frameworks. Both approaches recovered highly congruent topologies, consistently supporting a single, well-defined western Atlantic mitochondrial lineage with low internal divergence (0.04-2.13%). Species delimitation analyses based on multiple methods yielded partially congruent results but consistently identified a dominant lineage encompassing all Atlantic samples. In contrast, two Colombian reference mitogenomes formed a separate and highly divergent lineage relative to the Atlantic group, despite showing moderate divergence between them. Comparative mitogenomic analyses revealed conserved genome organization, nucleotide composition bias, codon usage, and transfer RNA (tRNA) structures. All PCGs evolved under strong purifying selection, with Ka/Ks ratios well below unity. CONCLUSIONS: These results support mitochondrial genetic continuity across the Western Atlantic Narcine populations and do not provide mitochondrial evidence for multiple evolutionary lineages within the Western Atlantic. The marked mitochondrial divergence of Colombian reference mitogenomes highlights potential issues in sequence attribution and underscores the importance of data curation. Overall, complete mitochondrial genomes provide a robust framework for species delimitation and future integrative taxonomic assessments within Narcine.

Animals↗

One-time duplication and ongoing loss of mitochondrial tRNA genes in Cryptocercus cockroaches.

Mitochondrial genome is a popular marker in phylogenetics and species diversity estimations. Mitogenome is relatively compact and conserved, while gene rearrangements were found in some species across various organisms. Models to explain the origin and evolution of gene rearrangement have been proposed but seldom demonstrated; empirical evidence from closely related species is particularly scarce. Here, through an intensive case study of the cockroach genus Cryptocercus Scudder, 1862, we elucidate the evolution of mitochondrial gene order. This study utilized 51 new samples and re-assembled raw reads of 26 published samples. A diversity of rearrangement patterns is recovered, especially in the tRNA gene cluster between ND3 and ND5, which is effectively explained by the duplication - random loss model. Specifically, the entire tRNA gene cluster was duplicated; this duplication is potentially facilitated by chance binding between the 3' end of ND5 gene and the ND3-trnA region during DNA replication. Furthermore, we reveal that one of the gene copies degenerated stochastically across lineages, directly contributing to the observed diversity in gene arrangement. Gene rearrangement patterns are apomorphies for certain clades, providing additional evidence for the inferred phylogeny and serving as potential indicators of species. This study underscores the importance of intensive sampling and rigorous data curation for deciphering the evolutionary mechanisms.

Duplication–random loss model↗

DXYS267: DYS393 and its X chromosome counterpart.

The GATA repeat DYS393 was reported in 1987 among other Y-specific short tandem repeats. It has since been used for forensic and evolutionary studies. We decided to test its Y-specificity when we found that female DNA gave amplicons, in agreement with recent GDB-recorded experiences on radiation hybrids. Parent-child triplets revealed that heterozygous daughters always carried the same paternally derived amplicon which, however, was not amplified in their fathers' DNAs. The X-assignment was verified in larger families. A half-new primer set with a new reverse DYS393 primer, outside the old one, resulted in X amplicons in females as well as Y and X amplicons in males. This new primer set defines the new DXYS267 (GDB Data Curation). DNA-sequencing revealed four base pair differences between the Y- and the X-sequences. Two are within the reverse primer site sequence, thus probably causing preferential hybridization to the Y sequence when using the conventional primers. The two others are within the repeat array, giving the regular repeat GATA in the Y-sequence, and TATA and GACA, respectively, in the X-sequence. Allele frequency distribution in DYS393 was studied in 300 unrelated Norwegian males, allele distribution in the X-locus in 48 Norwegian women. Even if allele repeat numbers are overlapping between the loci, leading to identical fragment lengths, the allele distribution is different between DYS393 and the X-chromosome locus. The differences between the two homologous loci on the Y and X indicate a considerable lap of time since common ancestry. To avoid co-amplification of the X-locus in DYS393 typing, primer A was elongated to include one of the sequence differences between the two loci. This to a considerable extent improved the specificity of the DYS393 primers.

Adenine↗

Word sense disambiguation in the biomedical domain: an overview.

There is a trend towards automatic analysis of large amounts of literature in the biomedical domain. However, this can be effective only if the ambiguity in natural language is resolved. In this paper, the current state of research in word sense disambiguation (WSD) is reviewed. Several methods for WSD have already been proposed, but many systems have been tested only on evaluation sets of limited size. There are currently only very few applications of WSD in the biomedical domain. The current direction of research points towards statistically based algorithms that use existing curated data and can be applied to large sets of biomedical literature. There is a need for manually tagged evaluation sets to test WSD algorithms in the biomedical domain. WSD algorithms should preferably be able to take into account both known and unknown senses of a word. Without WSD, automatic metaanalysis of large corpora of text will be error prone.

Algorithms↗

Weighting hidden Markov models for maximum discrimination.

MOTIVATION: Hidden Markov models can efficiently and automatically build statistical representations of related sequences. Unfortunately, training sets are frequently biased toward one subgroup of sequences, leading to an insufficiently general model. This work evaluates sequence weighting methods based on the maximum-discrimination idea. RESULTS: One good method scales sequence weights by an exponential that ranges between 0.1 for the best scoring sequence and 1.0 for the worst. Experiments with a curated data set show that while training with one or two sequences performed worse than single-sequence Probabilistic Smith-Waterman, training with five or ten sequences reduced errors by 20% and 51%, respectively. This new version of the SAM HMM suite outperforms HMMer (17% reduction over PSW for 10 training sequences), Meta-MEME (28% reduction), and unweighted SAM (31% reduction). AVAILABILITY: A WWW server, as well as information on obtaining the Sequence Alignment and Modeling (SAM) software suite and additional data from this work, can be found at http://www.cse.ucse. edu/research/compbio/sam.html

Algorithms↗

Improved prediction of the number of residue contacts in proteins by recurrent neural networks.

Knowing the number of residue contacts in a protein is crucial for deriving constraints useful in modeling protein folding, protein structure, and/or scoring remote homology searches. Here we use an ensemble of bi-directional recurrent neural network architectures and evolutionary information to improve the state-of-the-art in contact prediction using a large corpus of curated data. The ensemble is used to discriminate between two different states of residue contacts, characterized by a contact number higher or lower than the average value of the residue distribution. The ensemble achieves performances ranging from 70.1% to 73.1% depending on the radius adopted to discriminate contacts (6Ato 12A). These performances represent gains of 15% to 20% over the base line statistical predictors always assigning an aminoacid to the most numerous state, 3% to 7% better than any previous method. Combination of different radius predictors further improves the performance. SERVER: http://promoter.ics.uci.edu/BRNN-PRED/.

Amino Acid Sequence↗

NodMutDB: a database for genes and mutants involved in symbiosis.

UNLABELLED: Functional genomics research is producing large amounts of data on the functions of individual genes related to symbiosis. We have developed a relational database, NodMutDB (Nodulation Mutant Database), to provide a comprehensive resource for depositing, organizing and retrieving information on symbiosis-related genes, mutants and published literature. NodMutDB brings together new studies and existing mutant-based literature to facilitate our understanding of how genes function in symbiotic processes in both Rhizobia and their host plants. AVAILABILITY: http://nodmutdb.vbi.vt.edu CONTACT: cmao@vbi.vt.edu SUPPLEMENTARY INFORMATION: Database schema and data curation model are available at http://nodmutdb.vbi.vt.edu.

Bacterial Proteins↗

NQ-Flipper: validation and correction of asparagine/glutamine amide rotamers in protein crystal structures.

The error rate of asparagine (Asn) and glutamine (Gln) amide rotamers in protein crystal structures is in the order of 20% and as a consequence the current Protein Database (PDB) contains approximately half a million incorrect Asn and Gln side-chain rotamers. Here we present NQ-Flipper, a web service based on knowledge-based potentials of mean force to automatically detect and correct erroneous rotamers. We achieve excellent agreement with expert curated data.

Asparagine↗

DNA surveillance: web-based molecular identification of whales, dolphins, and porpoises.

DNA Surveillance is a Web-based application that assists in the identification of the species and population of unknown specimens by aligning user-submitted DNA sequences with a validated and curated data set of reference sequences. Phylogenetic analyses are performed and results are returned in tree and table format summarizing the evolutionary distances between the query and reference sequences. DNA Surveillance is implemented with mitochondrial DNA (mtDNA) control region sequences representing the majority of recognized cetacean species. Extensions of the system to include other gene loci and taxa are planned. The service, including instructions and sample data, is available at http://www.dna-surveillance.auckland.ac.nz.

Animals↗

WormBase: network access to the genome and biology of Caenorhabditis elegans.

WormBase (http://www.wormbase.org) is a web-based resource for the Caenorhabditis elegans genome and its biology. It builds upon the existing ACeDB database of the C.elegans genome by providing data curation services, a significantly expanded range of subject areas and a user-friendly front end.

Alternative Splicing↗

GrainGenes, the genome database for small-grain crops.

GrainGenes, http://www.graingenes.org, is the international database for the wheat, barley, rye and oat genomes. For these species it is the primary repository for information about genetic maps, mapping probes and primers, genes, alleles and QTLs. Documentation includes such data as primer sequences, polymorphism descriptions, genotype and trait scoring data, experimental protocols used, and photographs of marker polymorphisms, disease symptoms and mutant phenotypes. These data, curated with the help of many members of the research community, are integrated with sequence and bibliographic records selected from external databases and results of BLAST searches of the ESTs. Records are linked to corresponding records in other important databases, e.g. Gramene's EST homologies to rice BAC/PACs, TIGR's Gene Indices and GenBank. In addition to this information within the GrainGenes database itself, the GrainGenes homepage at http://wheat.pw.usda.gov provides many other community resources including publications (the annual newsletters for wheat, barley and oat, monographs and articles), individual datasets (mapping and QTL studies, polymorphism surveys, variety performance evaluations), specialized databases (Triticeae repeat sequences, EST unigene sets) and pages to facilitate coordination of cooperative research efforts in specific areas such as SNP development, EST-SSRs and taxonomy. The goal is to serve as a central point for obtaining and contributing information about the genetics and biology of these cereal crops.

Alleles↗

Large scale study of protein domain distribution in the context of alternative splicing.

Alternative splicing plays an important role in processes such as development, differentiation and cancer. With the recent increase in the estimates of the number of human genes that undergo alternative splicing from 5 to 35-59%, it is becoming critical to develop a better understanding of its functional consequences and regulatory mechanisms. We conducted a large scale study of the distribution of protein domains in a curated data set of several thousand genes and identified protein domains disproportionately distributed among alternatively spliced genes. We also identified a number of protein domains that tend to be spliced out. Both the proteins having the disproportionately distributed domains as well as those with spliced-out domains are predominantly involved in the processes of cell communication, signaling, development and apoptosis. These proteins function mostly as enzymes, signal transducers and receptors. Somewhat surprisingly, 28% of all occurrences of spliced-out domains are not effected by straightforward exclusion of exons coding for the domains but by inclusion or exclusion of other exons to shift the reading frame while retaining the exons coding for the domains in the final transcripts.

Alternative Splicing↗

IntEnz, the integrated relational enzyme database.

IntEnz is the name for the Integrated relational Enzyme database and is the official version of the Enzyme Nomenclature. The Enzyme Nomenclature comprises recommendations of the Nomenclature Committee of the International Union of Bio chemistry and Molecular Biology (NC-IUBMB) on the nomenclature and classification of enzyme-catalysed reactions. IntEnz is supported by NC-IUBMB and contains enzyme data curated and approved by this committee. The database IntEnz is available at http://www.ebi.ac.uk/intenz.

Animals↗