Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Identification of a PAX-FKHR gene expression signature that defines molecular classes and determines the prognosis of alveolar rhabdomyosarcomas.

Alveolar rhabdomyosarcomas (ARMS) are aggressive soft-tissue sarcomas affecting children and young adults. Most ARMS tumors express the PAX3-FKHR or PAX7-FKHR (PAX-FKHR) fusion genes resulting from the t(2;13) or t(1;13) chromosomal translocations, respectively. However, up to 25% of ARMS tumors are fusion negative, making it unclear whether ARMS represent a single disease or multiple clinical and biological entities with a common phenotype. To test to what extent PAX-FKHR determine class and behavior of ARMS, we used oligonucleotide microarray expression profiling on 139 primary rhabdomyosarcoma tumors and an in vitro model. We found that ARMS tumors expressing either PAX-FKHR gene share a common expression profile distinct from fusion-negative ARMS and from the other rhabdomyosarcoma variants. We also observed that PAX-FKHR expression above a minimum level is necessary for the detection of this expression profile. Using an ectopic PAX3-FKHR and PAX7-FKHR expression model, we identified an expression signature regulated by PAX-FKHR that is specific to PAX-FKHR-positive ARMS tumors. Data mining for functional annotations of signature genes suggested a role for PAX-FKHR in regulating ARMS proliferation and differentiation. Cox regression modeling identified a subset of genes within the PAX-FKHR expression signature that segregated ARMS patients into three risk groups with 5-year overall survival estimates of 7%, 48%, and 93%. These prognostic classes were independent of conventional clinical risk factors. Our results show that PAX-FKHR dictate a specific expression signature that helps define the molecular phenotype of PAX-FKHR-positive ARMS tumors and, because it is linked with disease outcome in ARMS patients, determine tumor behavior.

Biomarkers, Tumor↗

Genomics and proteomics of bone cancer.

Although the control of bone metastasis has been the focus of intensive investigation, relatively little is known about the molecular mechanisms that regulate or predict the process, even though widespread skeletal dissemination is an important step in the progression of many tumors. As a result, understanding the complex interactions contributing to the metastatic behavior of tumor cells is essential for the development of effective therapies. Using a state-of-the-art combination of gene expression profiling and functional annotation of human tumor cells, and surface-enhanced laser desorption/ionization time-of-flight mass spectrometry of patient serum, we have shown that changes in tumor biochemistry correlate with disease progression and help to define the aggressive tumor phenotype. Based on these approaches, it is apparent that the metastatic phenotype of tumor cells is extremely complex. The identification of the phenotype of tumor cells has benefited greatly from the application of gene expression profiling (microarray analysis). This technology has been used by many investigators to identify changes in gene expression and cytokine and growth factor elaboration (such as interleukin 8). The tumor phenotype(s) presumably also include changes in the cell surface carbohydrate profile (via altered glycosyltransferase expression) and heparan sulfate expression (via increased heparanase activity), to name but a few. These specific alterations in gene expression, identified by functional annotation of accumulated microarray data, have been validated using a variety of approaches. Collectively, the data described here suggest that each of these activities is associated with distinct aspects of the aggressive tumor cell phenotype. Collectively, the data suggest that multiple factors constitute the complex phenotype of metastatic tumor cells. In particular, the differences observed in gene expression profiles and serum protein biomarkers play a critical role in defining the mechanisms responsible for bone-specific colonization and growth of tumors in bone. Future studies will identify the mechanisms that participate in the formation of secondary tumor growths of cancers in bone.

Biomarkers, Tumor↗

Array-based comparative genomic hybridization and copy number variation in cancer research.

Array-based comparative genomic hybridization (aCGH) is a molecular cytogenetic technique used in detecting and mapping DNA copy number alterations. aCGH is able to interrogate the entire genome at a previously unattainable, high resolution and has directly led to the recent appreciation of a novel class of genomic variation: copy number variation (CNV) in mammalian genomes. All forms of DNA variation/polymorphism are important for studying the basis of phenotypic diversity among individuals. CNV research is still at its infancy, requiring careful collation and annotation of accumulating CNV data that will undoubtedly be useful for accurate interpretation of genomic imbalances identified during cancer research.

Animals↗

stam--a Bioconductor compliant R package for structured analysis of microarray data.

BACKGROUND: Genome wide microarray studies have the potential to unveil novel disease entities. Clinically homogeneous groups of patients can have diverse gene expression profiles. The definition of novel subclasses based on gene expression is a difficult problem not addressed systematically by currently available software tools. RESULTS: We present a computational tool for semi-supervised molecular disease entity detection. It automatically discovers molecular heterogeneities in phenotypically defined disease entities and suggests alternative molecular sub-entities of clinical phenotypes. This is done using both gene expression data and functional gene annotations. We provide stam, a Bioconductor compliant software package for the statistical programming environment R. We demonstrate that our tool detects gene expression patterns, which are characteristic for only a subset of patients from an established disease entity. We call such expression patterns molecular symptoms. Furthermore, stam finds novel sub-group stratifications of patients according to the absence or presence of molecular symptoms. CONCLUSION: Our software is easy to install and can be applied to a wide range of datasets. It provides the potential to reveal so far indistinguishable patient sub-groups of clinical relevance.

Calibration↗

De novo genome assemblies of threatened Asian hornbills (Bucerotidae) reveal declining population trajectories during the late Pleistocene.

BACKGROUND: Asian hornbills are flagship species of the wet tropics that face significant threats from hunting, habitat loss, and fragmentation. Despite being conservation flagships, whole genome information is available for only two of the 32 Asian hornbill species. In this study, we provide the first de novo genome assemblies for four hornbill species (Bucerotidae) in Asia. METHODS: We used a combination of long-read and short-read sequencing data to assemble and annotate de novo hybrid genomes of four species of hornbills. We also assembled and compared mitochondrial genomes of these species. Using a comparative genomics approach, we performed orthology assignment and gene evolution analyses to identify unique gene families in Asian hornbills, gene families that showed significant expansion, their functions and structural variation. Furthermore, using the Pairwise Sequentially Markov Coalescent (PSMC) method, we reconstructed demographic histories of hornbill species to examine changes in their population trajectories in the past. RESULTS: We present hybrid genome assemblies for Great Hornbill (B. bicornis - GH), Rufous-necked Hornbill (A. nipalensis- RNH), Malabar Pied Hornbill (A. coronatus- MPH) and Wreathed Hornbill (R. undulatus- WH). The genome sizes of these hornbills range from 1.1 Gb to 1.3 Gb, with over 95.9% completeness and gene prediction BUSCO. We reported 10,525 orthogroups shared among four Asian hornbill species and identified significant expansion in gene families associated with structural keratin development in Asian hornbills compared to their ancestors. We also provide annotated mitogenomes for each of these species. Furthermore, we found that the WH, a more abundant, widely distributed, and migratory species, showed a higher Ne than the other three hornbill species. However, an overall decline in Ne for all species was recorded during the Pleistocene climatic fluctuations. CONCLUSIONS: We present the first-ever, high-quality reference genomes for the threatened hornbill species from Asia. Hornbills have shown significant expansion in genes involved in structural keratin development. Our results indicate that Pleistocene climatic fluctuations have led to dramatic population declines in all four species. We believe that this study provides robust genomic resources to support future comparative and conservation genomics efforts for hornbills.

Animals↗

The Schistosoma mansoni gene index: gene discovery and biology by reconstruction and analysis of expressed gene sequences.

Expressed sequence tag (EST) sequencing and analysis is a primary research tool to identify and characterize the Schistosoma mansoni transcriptome. As part of our gene discovery effort, a total of 5,793 ESTs have been generated from clones selected randomly from complementary DNA (cDNA) libraries constructed from male and female adult worms. Assembly analysis of all the 16,813 public S. mansoni ESTs has identified 1,920 distinct tentative consensus sequences (TCs) and 5,571 nonoverlapping ESTs (singletons). Of these, 376 TCs (20%) and 1,449 singletons (26%) are unique to the SUNY/TIGR sequencing effort. Tentative consensus sequences and singletons were distributed into various categories of biological roles associated with cell structure, metabolism, protein fate, signal transduction, transcription, protein synthesis, transporters, and cell growth. The TCs and singletons represent transcripts that can be used as a resource for functional annotation of genomic sequence data, comparative sequence analysis, and cDNA clone selection for microarray projects. The utility of EST analysis is demonstrated by identifying new protease genes, which may be involved in hemoglobin degradation.

Amino Acid Sequence↗

The Human Genome Project: the role of analytical chemists.

The Human Genome Project (HGP) is the most ambitious and important effort in the history of biology. It has provided a complete genetic blueprint for human life, and will provide important insights into human health and development. HGP involves a huge amount of data that is stored on computers all over the world. More than just vast amounts of DNA sequences, the project is about developing sets of integrated maps that involve genetic, physical, and sequence data. The data can be sorted, annotated and organized in many different ways using different types of database software, different analysis algorithms and different forms of interfaces. The genomic sequences of the human and the substantial portions of the mouse genome are expected to be finished by 2005. Analytical chemists took the opportunity, addressing the problem of achieving a high throughput with good sensitivity. This paper discusses how analytical chemists saved the Human Genome Project or at least gave it a helping hand.

Chemistry Techniques, Analytical↗

Computer-assisted image analysis protocol that quantitatively measures subnuclear protein organization in cell populations.

Many nuclear proteins, including the nuclear receptor co-repressor (NCoR) protein are localized to specific regions of the cell nucleus, and this subnuclear positioning is preserved when NCoR is expressed in cells as a fusion to a fluorescent protein (FP). To determine how specific factors may influence the subnuclear organization of NCoR requires an unbiased approach to the selection of cells for image analysis. Here, we use the co-expression of the monomeric red FP (mRFP) to select cells that also express NCoR labeled with yellow FP (YFP). The transfected cells are selected for imaging based on the diffuse cellular mRFP signal without prior knowledge of the subnuclear organization of the co-expressed YFP-NCoR. The images acquired of the expressed FPs are then analyzed using an automated image analysis protocol that identifies regions of interest (ROIs) using a set of empirically determined rules. The relative expression levels of both fluorescent proteins are estimated, and YFP-NCoR subnuclear organization is quantified based on the mean focal body size and relative intensity. The selected ROIs are tagged with an identifier and annotated with the acquired data. This integrated image analysis protocol is an unbiased method for the precise and consistent measurement of thousands of ROIs from hundreds of individual cells in the population.

Animals↗

Aptamer-dependent full-length cDNA synthesis by overlap extension PCR.

Sequencing of the human genome in combination with computational annotation has provided tremendous data on predicted genes. However, for most of them, no corresponding cDNAs are available yet. Furthermore, even where cDNA clones were obtained, gene transcripts often have many different splice variants that are not covered by current gene collections. For direct synthesis of cDNA clones corresponding to predicted genes, new splice variants, or any other gene of interest, we established optimal PCR conditions for the direct amplification of exons from genomic DNA, which require a specific Taq aptamer. PCR products comprising differently tagged exons were concatenated by overlap extension into full-length cDNAs. To prove the effectiveness of the approach, the 1900-bp full-length open reading frame of the human mitochondrial aldehyde dehydrogenase (ALDH2) gene was synthesized in a two-step reaction comprising all 13 exons. Thus, our conditions are of general value for in vitro synthesis of cDNAs and alternative splice variants from genomic DNA.

Aldehyde Dehydrogenase↗

Hierarchical docking of databases of multiple ligand conformations.

Ligand flexibility is an important problem in molecular docking and virtual screening. To address this challenge, we investigate a hierarchical pre-organization of multiple conformations of small molecules. Such organization of pre-calculated conformations removes the exploration of ligand conformational space from the docking calculation and allows for concise representation of what can be thousands of conformations. The hierarchy also recognizes and prunes incompatible conformations early in the calculation, eliminating redundant calculations of fit. We investigate the method by docking the MDL Drug Data Report (MDDR), an annotated database of 100,000 molecules, into apo and holo forms of seven unrelated targets. This annotated database allows us to track the ranking of tens to hundreds of annotated ligands in each of the docking systems. The binding sites and database are prepared in an automated fashion in an attempt to remove some human bias from the calculations. Many thousands of explicit and implicit ligand conformations may be docked in calculations not much longer than required for single conformer docking. As long as internal energies are not considered, recombination with the hierarchy is additive as the number of degrees of freedom is increased. Molecules with even millions of conformations can be docked in a few minutes on a single desktop computer.

Computer Simulation↗

Comparative Genomic Analysis of Six Mycoplasma Gallisepticum Strains: Insights into Genetic Diversity and Antibiotic Resistance.

Mycoplasma gallisepticum (MG) is a significant pathogen that causes respiratory diseases, which have had a substantial economic impact on the poultry industry. Despite the resistance of MG to antibiotics, it is imperative to identify genetic diversity in order to develop countermeasures. In this study, the genomes of six MG strains were examined to gain deeper insights into the mutations. The data pertaining to Variant Annotation and Mutation Analysis using SnpEff, along with the calculation of mutation rates as the ratio of total mutations to the length of the genomic regions analyzed, were thoroughly examined. The comprehensive evaluation yielded a total of 25,942 variants across the six strains, underscoring substantial genetic diversity. Notably, strain S6 exhibited a preponderance of frameshift mutations. A notable finding was the presence of a mutation in the MsbA gene shared by all six strains. Furthermore, five of the six strains, with the exception of strain F99 Lab, exhibited a mutation at position 5158, which impacts a multidrug transport system. Notably, strain ATCC exhibits a distinctive mutation at position 942, while strain S6 displays a unique mutation at position 6855, which is linked to efflux ABC transporter components. Furthermore, a substantial degree of genetic variation was observed among the CrmA, GapA, and vlhA genes among the various strains. High-impact changes, such as insertions and deletions, exhibited a higher frequency in CrmA, particularly in strain S6. Conversely, nonsynonymous variations demonstrated a heightened prevalence in GapA, particularly in strain F99 Lab. The vlhA gene exhibited a spectrum of effects, ranging from synonymous mutations to high-impact mutations such as stop-gains and frameshifts, particularly in strains k5111a and k4602. The functional variations observed among the strains can be attributed to these mutations, which have the potential to alter gene expression or protein function. Furthermore, substantial mutations in the dxr and rpoC genes were associated with antibiotic resistance. These mutations underscore the ongoing evolutionary adaptations of M. gallisepticum. Consequently, there is an imperative for the revision of treatment protocols and the formulation of targeted vaccines to regulate resistance within the poultry industry.

Mycoplasma gallisepticum↗

Insights into phylogenetic relationships of Veronica species (Plantaginaceae) based on comparative chloroplast genomics.

INTRODUCTION: Veronica L. is one of the most species-rich genera in Plantaginaceae and several species have medicinal, horticultural, or ecological value. METHODS: In this study, the complete chloroplast genomes of three Veronica species were assembled and annotated using Illumina sequencing data. RESULTS: The plastomes exhibited a typical quadripartite structures, with total lengths of 150,202 bp for Veronica biloba L., 151,159 bp for Veronica ciliata Fisch. and 151,098 bp for Veronica vandellioides Maxim. Each genome contained 130-132 unique genes, including 86-87 protein-coding genes, 36-37 tRNA genes, and 8 rRNA genes. Comparative analyses of 24 Veronica plastomes indicated that the IR/SC junctions were largely conserved, although slight boundary shifts occurred around rps19, ndhF, and ycf1. Forward, palindromic, complement, and reverse repeats were detected, and A/T mononucleotide repeats were the dominant SSR type. Nucleotide diversity analysis identified rpl32-trnL, trnK-rps16, rpl32, ycf1, ndhF, accD, matK, and rpoB as highly variable regions. Phylogenetic analyses recovered Veronica as a well-supported monophyletic lineage and clarified the plastid positions of the three newly sequenced species. Divergence time estimation suggested that the estimation suggested of Veronica was around 14.9 Ma, with V. biloba, V. ciliata and V. vandellioides diverging approximately 3.9 Ma, 0.6 Ma, and 6.9 Ma, respectively. DISCUSSION: Because the analyses were based on plastid genomes, the inferred topology should be interpreted as chloroplast phylogenetic evidence rather than a complete species-history reconstruction. These results provide plastome resources and molecular evidence for taxonomy, species identification, and future evolutionary studies of Veronica.

Plantaginaceae↗

Identification of novel highly expressed genes in pancreatic ductal adenocarcinomas through a bioinformatics analysis of expressed sequence tags.

In most microarray experiments, a significant fraction of the differentially expressed mRNAs identified correspond to expressed sequence tags (ESTs) and are generally discarded from further analyses. We used careful bioinformatics analyses to characterize those ESTs that were found to be highly overexpressed in a series of pancreatic adenocarcinomas. cDNA was prepared from 60 non-neoplastic samples (normal pancreas [n = 20], normal colon [n = 10], or normal duodenal mucosal [n = 30]) and from 64 pancreatic cancers (resected cancers [n = 50] or cancer cell lines [n = 14]) and hybridized to the complete Affymetrix Human Genome U133 GeneChip(R) set (arrays U133A and B) for simultaneous analysis of 45,000 fragments corresponding to 33,000 known genes and 6,000 ESTs. The GeneExpress(R) software system Fold Change Analysis Tool was used and 60 ESTs were identified that were expressed at levels at least 3-fold greater in the pancreatic cancers as compared to normal tissues. Searches against the human genomic sequence and comparative genomic analysis of human and mouse genomes was carried out using basic local alignment search tools (BLAST), BLASTN, and BLASTX, for identifying protein coding genes corresponding to the ESTs. Subsequently, in order to pick the most relevant candidate genes for a more detailed analysis, we looked for domains/motifs in the open reading frames using SMART and Pfam programs. We were able to definitively map 43 of the 60 ESTs to known or novel genes, and 15 of the ESTs could be localized in close proximity to a gene in the human genome although we were unable to establish that the EST was indeed derived from those genes. The differential expression of a subset of genes was confirmed at the protein level by immunohistochemical labeling of tissue microarrays (inhibin beta A [INHBA] and CD29) and/or at the transcript level by RT-PCR (INHBA, AKAP12, ELK3, FOXQ1, EIF5A2, and EFNA5). We conclude that bioinformatics tools can be used to characterize differentially overexpressed ESTs, and that some of these ESTs may represent diagnostically and therapeutically useful targets that might be missed using data solely from currently annotated databases.

Adenocarcinoma↗

[Expression and catalysis of glucokinase of Thermoanaerobacter tengcongensis at different temperatures].

According to analysis of proteomic profiling for Thermoanaerobacter tencongensis, TTE0090 could be a novel gene of glucokianse (GLK), though no GLK gene was annotated in the genomic data. With the methods of cloning and expression in vitro, the recombinant TTE0090 was successfully expressed and purified. The recombinant TTE0090 exhibited the catalysis of GLK, even at high temperatures. Detection of expression levels and catalysis of TTE0090 in vivo was furthermore carried out at different temperatures. The expression of TTE0090 was attenuated during the culture temperature elevated; however, the specific activity was positively correlated to temperature raised. This leads a possibility that the metabolic capacity of glycolysis in T. tencongensis is relatively constant at different temperatures. All the results herein demonstrate that TTE0090 is a novel gene of GLK. The studies on TTE0090 and its protein product, thus, may deepen our understanding of the adaptation mechanism of thermophilic bacteria living in harsh environment.

Bacterial Proteins↗

Immunoreactive outer membrane proteins of Leptospira interrogans serovar Canicola strain Hond Utrecht IV.

BACKGROUND & OBJECTIVES: Leptospirosis is a severe and complex zoonotic disease prevalent in many countries including India. Current leptospiral research is focussed on the identification of the outer membrane proteins (OMPs) of the organism that could be used in developing diagnostic assays for leptospirosis. METHODS: The Leptospira interrogans serovar Canicola was grown in EMJH medium and the cells were subjected to sarcosyl detergent treatment. The sarcosyl soluble (SS) and sarcosyl insoluble (SI) fractions were analyzed by SDS-PAGE and immunoblotting to deduce their protein profile and identifying various immunodominant antigens. RESULTS: The protein profile of SS fractions indicated the presence of three major bands of 41, 32 and 25 kDa and minor bands of 85 and 46 kDa. The SI fraction in serovar Canicola revealed the presence of 112, 93, 77, 43, 36, 29 and 22.5 kDa as major bands and minor bands of 102 and 53 kDa. In immunoblotting, the SS proteins of 41, 32 and 25 kDa and SI proteins of 112, 77, 36 and 22.5 kDa were detected to be major immunogenic proteins. INTERPRETATION & CONCLUSION: In our study immunogenic proteins were extracted from SS and SI fractions and OMPs were similar to those reported in other pathogenic Leptospira strains. These OMPs being unique to all the pathogenic leptospires, can be targeted for diagnostic purpose. Further analysis of the cellular location and expression of leptospiral proteins will be useful in the annotation of genomic sequence data and in providing insight into the biology of Leptospira cells.

Animals↗

[Validation of a discontinuously recording, digital long-term ECG system (Siemens-Sirecust 802/850) using a single beat analysis].

Using beat-to-beat analysis we studied the annotation of 44 data-bases of the Massachusetts Institute of Technology (MIT) in comparison to the classification performed by a new microprocessor of a 24-h-ambulatory electrocardiographic device, which is based on real-time analysis and a solid-state ECG documentation. QRS detection was performed with an accuracy of 99%. Sensitivity and positive predictive accuracy were 70% and 87% for supraventricular ectopy. Using a fixed prematurity index of 80%, ventricular ectopy with a total of 7,845 beats was identified with a sensitivity of 64% and a positive predictive accuracy of 97%. With the additional consideration of late premature ventricular beats (PVB) sensitivity increased to 91% with a positive predictive accuracy of 96%. A sensitivity of more than 80% for singular PVB was achieved in 28/33 databases (85%), a similar positive accuracy was achieved in 27/33 databases (82%). Altogether, ventricular pairs and ventricular tachycardia resulted in a sensitivity of 83% and 76%, respectively, and in a positive predictive accuracy of 87% and 73%, respectively. Sensitivity exceeded 80% for ventricular pairs in 11/15 databases (77%) and for ventricular tachycardia in 10/14 databases (71%); similar results were observed for the positive predictive accuracy with 11/15 (73%) and 9/14 (64%) databases. In 42/44 databases and in all databases for arrhythmias of Lown class IVA and IVB, Lown classification was determined correctly. The Sirecust-Holter-ECG-system, a new device with real-time analysis and solid-state memory results in an accuracy for singular and complex ventricular arrhythmias, comparable to some of the presently available Holter systems.

Arrhythmias, Cardiac↗

Application of a minicomputer-based system in measuring intraocular fluid dynamics.

A complete, computerized system has been developed to automate and display radionuclide clearance studies in an ophthalmology clinical laboratory. The system is based on a PDP-8E computer with a 16-k core memory and includes a dual-drive Decassette system and an interactive display terminal. The software controls the acquisition of data from an NIM scaler, times the procedures, and analyzes and simultaneously displays logarithmically converted data on a fully annotated graph. Animal studies and clinical experiments are presented to illustrate the nature of these displays and the results obtained using this automated eye physiometer.

Animals↗

Integrating large-scale genotype and phenotype data.

With the completion of the Human Genome Project, a new emphasis is focusing on the sequence variation and the resulting phenotype. The number of data available from genomic studies addressing this relationship is rapidly growing. In order to analyze these data as a whole, they need to be integrated, aggregated and annotated in a timely manner. The Pharmacogenetics and Pharmacogenomics Knowledge Base PharmGKB; ( ) assembles and disseminates these data and their associated metadata that are needed for unambiguous identification and replication. Assembling these data in a timely manner is challenging, and the scalability of these data produce major challenges for a knowledge base such as PharmGKB. However, it is only through rapid global meta-annotation of these data that we will understand the relationship between specific genotype(s) and the related phenotype. PharmGKB has confronted these challenges, and these experiences and solutions can benefit all genome communities.

Animals↗