Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,243 records · Page 69Linked to original sources

CAFTAN: a tool for fast mapping, and quality assessment of cDNAs.

BACKGROUND: The German cDNA Consortium has been cloning full length cDNAs and continued with their exploitation in protein localization experiments and cellular assays. However, the efficient use of large cDNA resources requires the development of strategies that are capable of a speedy selection of truly useful cDNAs from biological and experimental noise. To this end we have developed a new high-throughput analysis tool, CAFTAN, which simplifies these efforts and thus fills the gap between large-scale cDNA collections and their systematic annotation and application in functional genomics. RESULTS: CAFTAN is built around the mapping of cDNAs to the genome assembly, and the subsequent analysis of their genomic context. It uses sequence features like the presence and type of PolyA signals, inner and flanking repeats, the GC-content, splice site types, etc. All these features are evaluated in individual tests and classify cDNAs according to their sequence quality and likelihood to have been generated from fully processed mRNAs. Additionally, CAFTAN compares the coordinates of mapped cDNAs with the genomic coordinates of reference sets from public available resources (e.g., VEGA, ENSEMBL). This provides detailed information about overlapping exons and the structural classification of cDNAs with respect to the reference set of splice variants. The evaluation of CAFTAN showed that is able to correctly classify more than 85% of 5950 selected "known protein-coding" VEGA cDNAs as high quality multi- or single-exon. It identified as good 80.6 % of the single exon cDNAs and 85 % of the multiple exon cDNAs. The program is written in Perl and in a modular way, allowing the adoption of this strategy to other tasks like EST-annotation, or to extend it by adding new classification rules and new organism databases as they become available. We think that it is a very useful program for the annotation and research of unfinished genomes. CONCLUSION: CAFTAN is a high-throughput sequence analysis tool, which performs a fast and reliable quality prediction of cDNAs. Several thousands of cDNAs can be analyzed in a short time, giving the curator/scientist a first quick overview about the quality and the already existing annotation of a set of cDNAs. It supports the rejection of low quality cDNAs and helps in the selection of likely novel splice variants, and/or completely novel transcripts for new experiments.

Chromosome Mapping↗

A comparative method for identification of gene structures and alternatively spliced variants.

MOTIVATION: Alternative splicing (AS) serves as a mechanism to create diversity among functional proteins. Increasing evidence indicates that a large portion of genes have AS forms. Hence AS variants should be considered while analyzing gene structures. RESULTS: A new cross-species gene identification and AS analysis system, PSEP, has been developed. The system is based on expressed sequence tag (EST)-to-genome and genome-to-genome comparisons and is implemented in two steps: sequence alignment and a series of post-alignment processes, including progressive signal extraction and patching. For gene identification, these post-alignment processes serve as noise filters and enable PSEP to eliminate approximately 88% of potential overprediction. The overall accuracy of PSEP is better than or comparable to that of other well-known cross-species gene prediction programs, including the ROSETTA program, TWINSCAN, SGP-1/-2 and SLAM, when tested on three benchmark datasets (the ELN gene region, the HoxA cluster and the ROSETTA set). In addition, 76.2 and 76.0% of multiple-exon genes in the ROSETTA dataset and human chromosome 20, respectively, are found to have AS forms. Approximately 23% of the 210 elementary alternatives identified in the ROSETTA dataset are not conserved between the human and mouse genomes, and none of the 210 transcripts is found in the RefSeq annotation. With its dual functions in cross-species conserved sequence analysis and AS analysis, PSEP is highly suitable for studying the evolution of AS patterns and for finding unidentified gene expression features.

Algorithms↗

An object model and database for functional genomics.

MOTIVATION: Large-scale functional genomics analysis is now feasible and presents significant challenges in data analysis, storage and querying. Data standards are required to enable the development of public data repositories and to improve data sharing. There is an established data format for microarrays (microarray gene expression markup language, MAGE-ML) and a draft standard for proteomics (PEDRo). We believe that all types of functional genomics experiments should be annotated in a consistent manner, and we hope to open up new ways of comparing multiple datasets used in functional genomics. RESULTS: We have created a functional genomics experiment object model (FGE-OM), developed from the microarray model, MAGE-OM and two models for proteomics, PEDRo and our own model (Gla-PSI-Glasgow Proposal for the Proteomics Standards Initiative). FGE-OM comprises three namespaces representing (i) the parts of the model common to all functional genomics experiments; (ii) microarray-specific components; and (iii) proteomics-specific components. We believe that FGE-OM should initiate discussion about the contents and structure of the next version of MAGE and the future of proteomics standards. A prototype database called RNA And Protein Abundance Database (RAPAD), based on FGE-OM, has been implemented and populated with data from microbial pathogenesis. AVAILABILITY: FGE-OM and the RAPAD schema are available from http://www.gusdb.org/fge.html, along with a set of more detailed diagrams. RAPAD can be accessed by registration at the site.

Abstracting and Indexing↗

Streamlining large-scale genomic data management: Insights from the UK Biobank whole-genome sequencing data.

Biobank-scale whole-genome sequencing (WGS) studies are increasingly pivotal in unraveling the genetic bases of diverse health outcomes. However, managing and analyzing these datasets' sheer volume and complexity presents significant challenges. We highlight the annotated genomic data structure (aGDS) format, substantially reducing the WGS data file size while enabling seamless integration of genomic and functional information for comprehensive WGS analyses. The aGDS format yielded 23 chromosome-specific files for the UK Biobank 500k WGS dataset, occupying only 1.10 tebibytes of storage. We develop the vcf2agds toolkit that streamlines the conversion of WGS data from VCF to aGDS format. Additionally, the STAARpipeline equipped with the aGDS files enabled scalable, comprehensive, and functionally informed WGS analysis, facilitating the detection of common and rare coding and noncoding phenotype-genotype associations. Overall, the vcf2agds toolkit and STAARpipeline provide a streamlined solution that facilitates efficient data management and analysis of biobank-scale WGS data across hundreds of thousands of samples.

Humans↗

Genetic mapping of new cotton fiber loci using EST-derived microsatellites in an interspecific recombinant inbred line cotton population.

There is an immediate need for a high-density genetic map of cotton anchored with fiber genes to facilitate marker-assisted selection (MAS) for improved fiber traits. With this goal in mind, genetic mapping with a new set of microsatellite markers [comprising both simple (SSR) and complex (CSR) sequence repeat markers] was performed on 183 recombinant inbred lines (RILs) developed from the progeny of the interspecific cross Gossypium hirsutum L. cv. TM1 x Gossypium barbadense L. Pima 3-79. Microsatellite markers were developed using 1557 ESTs-containing SSRs (> or = 10 bp) and 5794 EST-containing CSRs (> or = 12 bp) obtained from approximately 14,000 consensus sequences derived from fiber ESTs generated from the cultivated diploid species Gossypium arboreum L. cv AKA8401. From a total of 1232 EST-derived SSR (MUSS) and CSR (MUCS) primer-pairs, 1019 (83%) successfully amplified PCR products from a survey panel of six Gossypium species; 202 (19.8%) were polymorphic between the G. hirsutum L. and G. barbadense L. parents of the interspecific mapping population. Among these polymorphic markers, only 86 (42.6%) showed significant sequence homology to annotated genes with known function. The chromosomal locations of 36 microsatellites were associated with 14 chromosomes and/or 13 chromosome arms of the cotton genome by hypoaneuploid deficiency analysis, enabling us to assign genetic linkage groups (LG) to specific chromosomes. The resulting genetic map consists of 193 loci, including 121 new fiber loci not previously mapped. These fiber loci were mapped to 19 chromosomes and 11 LG spanning 1277 cM, providing approximately 27% genome coverage. Preliminary quantitative trait loci analysis suggested that chromosomes 2, 3, 15, and 18 may harbor genes for traits related to fiber quality. These new PCR-based microsatellite markers derived from cotton fiber ESTs will facilitate the development of a high-resolution integrated genetic map of cotton for structural and functional study of fiber genes and MAS of genes that enhance fiber quality.

Aneuploidy↗

The crystal structure of Rv0793, a hypothetical monooxygenase from M. tuberculosis.

Mycobacterium tuberculosis infects millions worldwide. The Structural Genomics Consortium for M. tuberculosis has targeted all genes from this bacterium in hopes of discovering and developing new therapeutic agents. Open reading frame Rv0793 from M. tuberculosis was annotated with an unknown function. The 3-dimensional structure of Rv0793 has been solved to 1.6 A resolution. Its structure is very similar to that of Streptomyces coelicolor ActVA-Orf6, a monooxygenase that participates in tailoring of polyketide antibiotics in the absence of a cofactor. It is also similar to the recently solved structure of YgiN, a quinol monooxygenase from Escherichia coli. In addition, the structure of Rv0793 is similar to several structures of other proteins with unknown function. These latter structures have been determined recently as a result of structural genomic projects for various bacterial species. In M. tuberculosis, Rv0793 and its homologs may represent a class of monooygenases acting as reactive oxygen species scavengers that are essential for evading host defenses. Since the most prevalent mode of attack by the host defense on M. tuberculosis is by reactive oxygen species and reactive nitrogen species, Rv0793 may provide a novel target to combat infection by M. tuberculosis.

Amino Acid Sequence↗

Genomes with distinct function composition.

The functional composition of organisms can be analysed for the first time with the appearance of complete or sizeable parts of various genomes. We have reduced the problem of protein function classification to a simple scheme with three classes of protein function: energy-, information- and communication-associated proteins. Finer classification schemes can be easily mapped to the above three classes. To deal with the vast amount of information, a system for automatic function classification using database annotations has been developed. The system is able to classify correctly about 80% of the query sequences with annotations. Using this system, we can analyse samples from the genomes of the most represented species in sequence databases and compare their genomic composition. The similarities and differences for different taxonomic groups are strikingly intuitive. Viruses have the highest proportion of proteins involved in the control and expression of genetic information. Bacteria have the highest proportion of their genes dedicated to the production of proteins associated with small molecule transformations and transport. Animals have a very large proportion of proteins associated with intra- and intercellular communication and other regulatory processes. In general, the proportion of communication-related proteins increases during evolution, indicating trends that led to the emergence of the eukaryotic cell and later the transition from unicellular to multicellular organisms.

Animals↗

A comprehensive search for HNF-3alpha-regulated genes in mouse hepatoma cells by 60K cDNA microarray and chromatin immunoprecipitation/PCR analysis.

To characterize the regulatory pattern by a specific transcription regulatory factor, we used a combination of expression analysis with the mouse cDNA microarray composed of 60,000 cDNA clones and cross-linking/chromatin immunoprecipitation (X-ChIP) followed by comparative PCR. Overexpression of mouse hepatocyte nuclear factor-3alpha (HNF-3alpha) in a mouse hepatoma cell line resulted in accompanied perturbed expression of more than 1500 genes. Search for HNF-3alpha consensus recognition sequences in the upstream regions of their coding sequences, which were mapped on the mouse genome, enabled us to mine 300 genes as the potential HNF-3alpha-regulated genes and classify 135 annotated ones into several functional categories. Further X-ChIP/PCR analysis demonstrated in vivo binding of HNF-3alpha to the 5(')-flanking sequences of 25 members selected out of these genes. Besides known HNF-3alpha-regulated genes such as albumin and alpha-fetoprotein genes, the genes newly identified as the HNF-3alpha-regulated ones include three encoding CDP-diacylglycerol-inositol 3-phosphatidyltransferase, phosphatidylserine decarboxylase, and phospholipase A2, which are located en suite in the lipid metabolic pathway in liver. The potential usefulness of the present approach to extensive characterization of gene expression framework directed by a specific transcription regulatory factor is discussed.

5' Flanking Region↗

Identification of the down-regulated genes in a mat1-2-deleted strain of Gibberella zeae, using cDNA subtraction and microarray analysis.

Gibberella zeae (anamorph: Fusarium graminearum), a self-fertile ascomycete, is an important pathogen of cereal crops. Here, we have focused on the genes specifically controlled by the mating type (MAT) locus, a master regulator of sexual developmental process in G. zeae. To identify these genes, we employed suppression subtractive hybridization between a G. zeae wild-type strain Z03643 and the isogenic self-sterile mat1-2 strain T43deltaM2-2. Both reverse Northern and cDNA microarray analyses using 291 subtractive unigenes confirmed that 58.8% (171 genes) were significantly down-regulated in T43deltaM2-2. Among these, 98 could be either manually or automatically annotated based on known functions of their possible homologs. Northern blot analysis revealed that all of the genes examined were differentially regulated by MAT1-2 during sexual development. This study is the first report on the set of genes that are transcriptionally altered by the deletion of MAT1-2 during sexual reproduction in G. zeae.

Blotting, Northern↗

Identification, cloning, and expression of bacteriophage T5 dnk gene encoding a broad specificity deoxyribonucleoside monophosphate kinase (EC 2.7.4.13).

The nucleotide sequence corresponding to 13-19.5% of the bacteriophage T5 genome in early region C was determined (GenBank AY 140897). One of the five major single-stranded interruptions (nicks) of bacteriophage T5 DNA was identified at 18.5%. The sequenced region was annotated and the putative functions of some open reading frames were proposed by comparison with databases. The dnk gene, encoding a deoxyribonucleoside monophosphate kinase, was identified using a previously defined N-terminal amino acid sequence. The gene was cloned and expressed in Escherichia coli, the enzyme was purified to homogeneity with high yield using two alternative methods, and the recombinant deoxyribonucleoside monophosphate kinase was found to have the same activity and specificity as the native enzyme.

Amino Acid Sequence↗

Telemedicine in practice.

Telemedicine is defined as the "delivery of health care and sharing of medical knowledge over a distance using telecommunication systems." The concept of telemedicine is not new. Beyond the use of the telephone, there were numerous attempts to develop telemedicine programs in the 1960s mostly based on interactive television. The early experience was conceptionally encouraging but suffered inadequate technology. With a few notable exceptions such as the telemetry of medical data in the space program, there was very little advancement of telemedicine in the 1970s and 1980s. Interest in telemedicine has exploded in the 1990s with the development of medical devices suited to capturing images and other data in digital electronic form and the development and installation of high speed, high bandwidth telecommunication systems around the world. Clinical applications of telemedicine are now found in virtually every specialty. Teleradiology is the most common application followed by cardiology, dermatology, psychiatry, emergency medicine, home health care, pathology, and oncology. The technological basis and the practical issues are highly variable from one clinical application to another. Teleradiology, including telenuclear medicine, is one of the more well-defined telemedicine services. Techniques have been developed for the acquisition and digitization of images, image compression, image transmission, and image interpretation. The American College of Radiology has promulgated standards for teleradiology, including the requirement for the use of high resolution 2000 x 2000 pixel workstations for the interpretation of plain films. Other elements of the standard address image annotation, patient confidentiality, workstation functionality, cathode ray tube brightness, and image compression. Teleradiology systems are now widely deployed in clinical practice. Applications include providing service from larger to smaller institutions, coverage of outpatient clinics, imaging centers, and nursing homes. Teleradiology is also being used in international applications. Unresolved issues in telemedicine include licensure, the development of standards, reimbursement for services, patient confidentiality, and telecommunications infrastructure and cost. A number of states and medical boards have instituted policies and regulations to prevent physicians who are not licensed in the respective state to provide telemedicine services. This is a major impediment to the delivery of telemedicine between states. Telemedicine, including teleradiology, is here to stay and is changing the practice of medicine dramatically. National and international communications networks are being created that enable the sharing of information and knowledge at a distance. Technological barriers are being overcome leaving organizational, legal, financial, and special interest issues as the major impediments to the further development of telemedicine and realization of its benefits.

Computer Communication Networks↗

Ego atrophy in substance abuse: addiction from a socio-cultural perspective.

The use of intoxicants is indexed in American history adopting a social perspective of the role of alcoholism in traditional American society. Appealing to societal patterns, the elaboration of substance abuse as a disease is explored with a diagnostic focus on intervention as it relates to pathogenesis. Using clinical vignettes, the ego is proposed as the focus of pathology in the addiction process, featuring a regression to the defenses of projection and denial. Following the model proposed in Zinberg's seminal paper on addiction and ego function, various deficits are annotated that typify such regression. The ultimate clinical picture is one of ego atrophy, where basic interests and object relations are usurped by the all-consuming preoccupation with the substance of abuse.

Behavior, Addictive↗

The complete genomes and proteomes of 27 Staphylococcus aureus bacteriophages.

Bacteriophages are the most abundant life forms in the biosphere. They play important roles in bacterial ecology, evolution, adaptation to new environments, and pathogenesis of human bacterial infections. Here, we report the complete genomic sequences, and predicted proteins of 27 bacteriophages of the Gram-positive bacterium Staphylococcus aureus. Comparative nucleotide and protein sequence analysis indicates that these phages are a remarkable source of untapped genetic diversity, encoding 2,170 predicted protein-encoding ORFs, of which 1,402 cannot be annotated for structure or function, and 522 are proteins with no similarity to other phage or bacterial sequences. Based on their genome size, organization of their gene map and comparative nucleotide and protein sequence analysis, the S. aureus phages can be organized into three groups. Comparison of their gene maps reveals extensive genome mosaicism, hinting to a large reservoir of unidentified S. aureus phage genes. Among the phages in the largest size class (178-214 kbp) that we characterized is phage Twort, the first discovered bacteriophage (responsible for the Twort-D'Herelle effect). These phage genomes offer an exciting opportunity to discern molecular mechanisms of phage evolution and diversity.

Chromosome Mapping↗

Molecular characterization of mouse gastric epithelial progenitor cells.

The adult mouse gastric epithelium undergoes continuous renewal in discrete anatomic units. Lineage tracing studies have previously disclosed the morphologic features of gastric epithelial lineage progenitors (GEPs), including those of the presumptive multipotent stem cell. However, their molecular features have not been defined. Here, we present the results of an analysis of genes and pathways expressed in these cells. One hundred forty-seven transcripts enriched in GEPs were identified using an approach that did not require physical disruption of the stem cell niche. Real-time quantitative RT-PCR studies of laser capture microdissected cells retrieved from this niche confirmed enriched expression of a selected set of genes from the GEP list. An algorithm that allows quantitative comparisons of the functional relatedness of automatically annotated expression profiles showed that the GEP profile is similar to a dataset of genes that defines mouse hematopoietic stem cells, and distinct from the profiles of two differentiated GEP descendant lineages (parietal and zymogenic cell). Overall, our analysis revealed that growth factor response pathways are prominent in GEPs, with insulin-like growth factor appearing to play a key role. A substantial fraction of GEP transcripts encode products required for mRNA processing and cytoplasmic localization, including numerous homologs of Drosophila genes (e.g., Y14, staufen, mago nashi) needed for axis formation during oogenesis. mRNA targeting proteins may help these epithelial progenitors establish differential communications with neighboring cells in their niche.

Adenosine Triphosphate↗

Mining the Giardia lamblia genome for new cyst wall proteins.

The Giardia lamblia cyst wall (CW), which is required for survival outside the host and infection, is a primitive extracellular matrix. Because of the importance of the CW, we queried the Giardia Genome Project Database with the coding sequences of the only two known CW proteins, which are cysteine-rich and contain leucine-rich repeats (LRRs). We identified five new LRR-containing proteins, of which only one (CWP3) is up-regulated during encystation and incorporated into the cyst wall. Sequence comparison with CWP1 and -2 revealed conservation within the LRRs and the 44-amino-acid N-flanking region, although CWP3 is more divergent. Interestingly, all 14 cysteine residues of CWP3 are positionally conserved with CWP1 and -2. During encystation, C-terminal epitope-tagged CWP3 was transported to the wall of water-resistant cysts via the novel regulated secretory pathway in encystation-secretory vesicles (ESVs). Deletion analysis revealed that the four LRRs are each essential to target CWP3 to the ESVs and cyst wall. In a deletion of the most C-terminal region, fewer ESVs were stained in encysting cells, and there was no staining in cysts. In contrast, deletion of the 44 amino acids between the signal sequence and the LRRs or the region just C-terminal to the LRRs only decreased the number of cells with CWP3 targeting to ESVs and cyst wall by approximately 50%. Our studies indicate that virtually every portion of the CWP3 protein is needed for efficient targeting to the regulated secretory pathway and incorporation into the cyst wall. Further, these data demonstrate the power of genomics in combination with rigorous functional analyses to verify annotation.

Amino Acid Sequence↗

Differential label-free quantitative proteomic analysis of Shewanella oneidensis cultured under aerobic and suboxic conditions by accurate mass and time tag approach.

We describe the application of LC-MS without the use of stable isotope labeling for differential quantitative proteomic analysis of whole cell lysates of Shewanella oneidensis MR-1 cultured under aerobic and suboxic conditions. LC-MS/MS was used to initially identify peptide sequences, and LC-FTICR was used to confirm these identifications as well as measure relative peptide abundances. 2343 peptides covering 668 proteins were identified with high confidence and quantified. Among these proteins, a subset of 56 changed significantly using statistical approaches such as statistical analysis of microarrays, whereas another subset of 56 that were annotated as performing housekeeping functions remained essentially unchanged in relative abundance. Numerous proteins involved in anaerobic energy metabolism exhibited up to a 10-fold increase in relative abundance when S. oneidensis was transitioned from aerobic to suboxic conditions.

Amino Acid Sequence↗

Database verification studies of SWISS-PROT and GenBank.

PROBLEM STATEMENT: We have studied the relationships among SWISS-PROT, TrEMBL, and GenBank with two goals. First is to determine whether users can reliably identify those proteins in SWISS-PROT whose functions were determined experimentally, as opposed to proteins whose functions were predicted computationally. If this information was present in reasonable quantities, it would allow researchers to decrease the propagation of incorrect function predictions during sequence annotation, and to assemble training sets for developing the next generation of sequence-analysis algorithms. Second is to assess the consistency between translated GenBank sequences and sequences in SWISS-PROT and TrEMBL. RESULTS: (1) Contrary to claims by the SWISS-PROT authors, we conclude that SWISS-PROT does not identify a significant number of experimentally characterized proteins. (2) SWISS-PROT is more incomplete than we expected in that version 38.0 from July 1999 lacks many proteins from the full genomes of important organisms that were sequenced years earlier. (3) Even if we combine SWISS-PROT and TrEMBL, some sequences from the full genomes are missing from the combined dataset. (4) In many cases, translated GenBank genes do not exactly match the corresponding SWISS-PROT sequences, for reasons that include missing or removed methionines, differing translation start positions, individual amino-acid differences, and inclusion of sequence data from multiple sequencing projects. For example, results show that for Escherichia coli, 80.6% of the proteins in the GenBank entry for the complete genome have identical sequence matches with SWISS-PROT/TrEMBL sequences, 13.4% have exact substring matches, and matches for 4.1% can be found using BLAST search; the remaining 2.0% of E.coli protein sequences (most of which are ORFs) have no clear matches to SWISS-PROT/TrEMBL. Although many of these differences can be explained by the complexity of the DB, and by the curation processes used to create it, the scale of the differences is notable.

Algorithms↗

A lock-and-key model for protein-protein interactions.

MOTIVATION: Protein-protein interaction networks are one of the major post-genomic data sources available to molecular biologists. They provide a comprehensive view of the global interaction structure of an organism's proteome, as well as detailed information on specific interactions. Here we suggest a physical model of protein interactions that can be used to extract additional information at an intermediate level: It enables us to identify proteins which share biological interaction motifs, and also to identify potentially missing or spurious interactions. RESULTS: Our new graph model explains observed interactions between proteins by an underlying interaction of complementary binding domains (lock-and-key model). This leads to a novel graph-theoretical algorithm to identify bipartite subgraphs within protein-protein interaction networks where the underlying data are taken from yeast two-hybrid experimental results. By testing on synthetic data, we demonstrate that under certain modelling assumptions, the algorithm will return correct domain information about each protein in the network. Tests on data from various model organisms show that the local and global patterns predicted by the model are indeed found in experimental data. Using functional and protein structure annotations, we show that bipartite subnetworks can be identified that correspond to biologically relevant interaction motifs. Some of these are novel and we discuss an example involving SH3 domains from the Saccharomyces cerevisiae interactome. AVAILABILITY: The algorithm (in Matlab format) is available (see http://www.maths.strath.ac.uk/~aas96106/lock_key.html).

Algorithms↗