Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Molecular Sequence Annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Evolution of the Rab family of small GTP-binding proteins.

Rab proteins are small GTP-binding proteins that form the largest family within the Ras superfamily. Rab proteins regulate vesicular trafficking pathways, behaving as membrane-associated molecular switches. Here, we have identified the complete Rab families in the Caenorhabditis elegans (29 members), Drosophila melanogaster (29), Homo sapiens (60) and Arabidopsis thaliana (57), and we defined criteria for annotation of this protein family in each organism. We studied sequence conservation patterns and observed that the RabF motifs and the RabSF regions previously described in mammalian Rabs are conserved across species. This is consistent with conserved recognition mechanisms by general regulators and specific effectors. We used phylogenetic analysis and other approaches to reconstruct the multiplication of the Rab family and observed that this family shows a strict phylogeny of function as opposed to a phylogeny of species. Furthermore, we observed that Rabs co-segregating in phylogenetic trees show a pattern of similar cellular localisation and/or function. Therefore, animal and fungi Rab proteins can be grouped in "Rab functional groups" according to their segregating patterns in phylogenetic trees. These functional groups reflect similarity of sequence, localisation and/or function, and may also represent shared ancestry. Rab functional groups can help the understanding of the functional evolution of the Rab family in particular and vesicular transport in general, and may be used to predict general functions for novel Rab sequences.

Amino Acid Motifs↗

LY-6K gene: a novel molecular marker for human breast cancer.

A full-length cDNA was identified using one STS sequence containing an SNP (single nucleotide polymorphism) derived from genomic DNAs of breast cancer patients using a variety of bioinformatics tools. The cDNA encodes LY-6K, a novel member protein of the Ly-6/uPAR superfamily. It has been annotated as a target antigen for the HNSCC (head-and neck squamous cell carcinoma). We isolated the LY-6K gene from genomic DNAs obtained from breast cancer patients through large scale, case-control-screening. We performed northern blot hybridization and semi-quantitative RT-PCR on a human multiple-tissue mRNA blot from several breast cancer patients. We investigated the expression level of the LY-6K gene in human breast cancer, and compared this to expression in human normal breast tissue. We found that LY-6K was more highly expressed in the mRNA of breast tumors compared to its expression in normal breast tissue. These results suggest that LY-6K is not only a target antigen for HNSCC but also a significant new molecular marker for diagnosis and gene therapy in patients with breast cancer.

Antigens, Ly↗

Rank information: a structure-independent measure of evolutionary trace quality that improves identification of protein functional sites.

Protein functional sites are key targets for drug design and protein engineering, but their large-scale experimental characterization remains difficult. The evolutionary trace (ET) is a computational approach to this problem that has been useful in a variety of case studies, but its proteomic scale application is partially hindered because automated retrieval of input sequences from databases often includes some with errors that degrade functional site identification. To recognize and purge these sequences, this study introduces a novel and structure-free measure of ET quality called rank information (RI). It is shown that RI decreases in response to errors in sequences, alignments, or functional classifications. Conversely, an automated procedure to increase RI by selectively removing sequences improves functional site identification so as to nearly match manually curated traces in kinases and in a test set of 79 diverse proteins. Thus we conclude that RI partially reflects the evolutionary consistency of sequence, structure, and function. In practice, as the size of the proteome continues to grow exponentially, it provides a novel and structure-free measure of ET quality that increases its accuracy for large-scale automated annotation of protein functional sites.

Algorithms↗

Proteome composition in Plasmodium falciparum: higher usage of GC-rich nonsynonymous codons in highly expressed genes.

The parasite Plasmodium falciparum, responsible for the most deadly form of human malaria, is one of the extremely AT-rich genomes sequenced so far and known to possess many atypical characteristics. Using multivariate statistical approaches, the present study analyzes the amino acid usage pattern in 5038 annotated protein-coding sequences in P. falciparum clone 3D7. The amino acid composition of individual proteins, though dominated by the directional mutational pressure, exhibits wide variation across the proteome. The Asn content, expression level, mean molecular weight, hydropathy, and aromaticity are found to be the major sources of variation in amino acid usage. At all stages of development, frequencies of residues encoded by GC-rich codons such as Gly, Ala, Arg, and Pro increase significantly in the products of the highly expressed genes. Investigation of nucleotide substitution patterns in P. falciparum and other Plasmodium species reveals that the nonsynonymous sites of highly expressed genes are more conserved than those of the lowly expressed ones, though for synonymous sites, the reverse is true. The highly expressed genes are, therefore, expected to be closer to their putative ancestral state in amino acid composition, and a plausible reason for their sequences being GC-rich at nonsynonymous codon positions could be that their ancestral state was less AT-biased. Negative correlation of the expression level of proteins with respective molecular weights supports the notion that P. falciparum, in spite of its intracellular parasitic lifestyle, follows the principle of cost minimization.

Amino Acids↗

Structure and action of urocanase.

Urocanase (EC 4.2.1.49) from Pseudomonas putida was crystallized after removing one of the seven free thiol groups. The crystal structure was solved by multiwavelength anomalous diffraction (MAD) using a seleno-methionine derivative and then refined at 1.14 A resolution. The enzyme is a symmetric homodimer of 2 x 557 amino acid residues with tightly bound NAD+ cofactors. Each subunit consists of a typical NAD-binding domain inserted into a larger core domain that forms the dimer interface. The core domain has a novel chain fold and accommodates the substrate urocanate in a surface depression. The NAD domain sits like a lid on the core domain depression and points with the nicotinamide group to the substrate. Substrate, nicotinamide and five water molecules are completely sequestered in a cavity. Most likely, one of these water molecules hydrates the substrate during catalysis. This cavity has to open for substrate passage, which probably means lifting the NAD domain. The observed atomic arrangement at the active center gives rise to a detailed proposal for the catalytic mechanism that is consistent with published chemical data. As expected, the variability of the residues involved is low, as derived from a family of 58 proteins annotated as urocanases in the data banks. However, one well-embedded member of this family showed a significant deviation at the active center indicating an incorrect annotation.

Amino Acid Sequence↗

NCBI genetic resources supporting immunogenetic research.

The NCBI creates and maintains a set of integrated bibliographic, sequence, map, structure and other database resources to promote the efficient retrieval of information and the discovery of novel relationships. The connections made between elements of these resources permit researchers to start a search from a wide spectrum of entry points. These multiple dimensions of data can be roughly categorized by primary content as text or bibliographic (PubMed, PubMedCentral, OMIM, LocusLink), sequence (GenBank, Reference Sequence Project (RefSeq), dbSNP, MMDB), protein structure (MMDB) or map position (MapView). They can also becategorized by level of expert curation, which may range from validation of submissions from external groups (GenBank, PubMed, PubMedCentral,), to automatic computation (HomoloGene, UniGene), and to highly reviewed and corrected (LocusLink, MMDB, OMIM, RefSeq). Searches can be made by words (in an article title, key words, sequence annotation, database value, author) by sequence (BLAST or e-PCR against multiple sequence databases), or by map coordinates. By computing or curating bi-directional links between related objects, NCBI can represent content on the genetics, molecular biology, and clinical considerations of interest to immunogeneticists. There is also an emerging resource developed by the NCBI in collaboration with the IHWG devoted to the presentation of MHC data (dbMHC). How dbMHC will augment existing resources at the NCBI is described.

Computational Biology↗

Annotation and BAC/PAC localization of nonredundant ESTs from drought-stressed seedlings of an indica rice.

To decipher the genes associated with drought stress response and to identify novel genes in rice, we utilized 1540 high-quality expressed sequence tags (ESTs) for functional annotation and mapping to rice genomic sequences. These ESTs were generated earlier by 3'-end single-pass sequencing of 2000 cDNA clones from normalized cDNA libraries constructed form drought-stressed seedlings of an indica rice. A rice UniGene set of 1025 transcripts was constructed from this collection through the BLASTN algorithm. Putative functions of 559 nonredundant ESTs were identified by BLAST similarity search against public databases. Putative functions were assigned at a stringency E value of 10(-6) in BLASTN and BLASTX algorithms. To understand the gene structure and function further, we have utilized the publicly available finished and unfinished rice BAC/PAC (BAC, bacterial artificial chromosome; PAC, P1 artificial chromosome) sequences for similarity search using the BLASTN algorithm. Further, 603 nonredundant ESTs have been mapped to BAC/PAC clones. BAC clones were assigned by a homology of above 95% identity along 90% of EST sequence length in the aligned region. In all, 700 ESTs showed rice EST hits in GenBank. Of the 325 novel ESTs, 128 were localized to BAC clones. In addition, 127 ESTs with identified putative functions but with no homology in IRGSP (International Rice Genome Sequencing Program) BAC/PAC sequences were mapped to the Chinese WGS (whole genome shotgun contigs) draft sequence of the rice genome. Functional annotation uncovered about a hundred candidate ESTs associated with abiotic stress in rice and Arabidopsis that were previously reported based on microarray analysis and other studies. This study is a major effort in identifying genes associated with drought stress response and will serve as a resource to rice geneticists and molecular biologists.

Chromosomes, Artificial, Bacterial↗

Mass spectrometric characterization of transferrins and their fragments derived by reduction of disulfide bonds.

Mass spectrometry, proteomics, and protein chemistry methods are used to characterize the cleavage products of 79 kDa transferrin proteins induced by iron-catalyzed oxidation, including a novel C-terminal polypeptide released upon disulfide reduction. Top-down electrospray ionization tandem mass spectrometry (ESI-MS/MS) of intact multiply-charged transferrin from a variety of species (human, bovine, rabbit, chicken) performed on a quadrupole time-of-flight mass spectrometer yields multiply-charged b(n)-products originating near residues 56-69 from the N-terminal region, in addition to their complementary y(n)-products. Incubation of transferrin with reductants, such as dithiothreitol (DTT) or tris(2-carboxyethyl)-phosphine (TCEP), yields an increase in multiple charging observed by ESI-MS and an increase in molecular weight consistent with disulfide reduction. However, mammalian transferrins release a 6-8 kDa fragment upon disulfide reduction. Protein acetylation and MS/MS sequencing demonstrate that the fragment originates from the C-terminus of the protein, and that it is a separate polypeptide linked via three disulfide bonds to the main transferrin chain. The existence of a separate C-terminal chain is not annotated in protein sequence databases and, to date, has not been reported in the literature. Iron-catalyzed cleavage induces fragments originating from both the N- and C-terminus of transferrin.

Amino Acid Sequence↗

Bioinformatics in support of molecular medicine.

Bioinformatics studies two important information flows in modern biology. The first is the flow of genetic information from the DNA of an individual organism up to the characteristics of a population of such organisms (with an eventual passage of information back to the genetic pool, as encoded within DNA). The second is the flow of experimental information from observed biological phenomena to models that explain them, and then to new experiments in order to test these models. The discipline of bioinformatics has its roots in a number of activities, including the organization of DNA sequence and protein three-dimensional structural data collections in the 1960's and 1970's. It has become a booming academic and industrial enterprise with the introduction of biological experiments that rapidly produce massive amounts of data (such as the multiple genome sequencing projects, the large scale analysis of gene expression, and the large scale analysis of protein-protein interactions). Basic biological science has always had an impact on clinical medicine (and clinical medical information systems), and is creating a new generation of epidemiologic, diagnostic, prognostic, and treatment modalities. Bioinformatics efforts that appear to be wholly geared towards basic science are likely to become relevant to clinical informatics in the coming decade. For example, DNA sequence information and sequence annotations will appear in the medical chart with increasing frequency. The algorithms developed for research in bioinformatics will soon become part of clinical information systems.

Computational Biology↗

Annotating eukaryote genomes.

The Genome Annotation Assessment Project tested current methods of gene identification, including a critical assessment of the accuracy of different methods. Two new databases have provided new resources for gene annotation: these are the InterPro database of protein domains and motifs, and the Gene Ontology database for terms that describe the molecular functions and biological roles of gene products. Efforts in genome annotation are most often based upon advances in computer systems that are specifically designed to deal with the tremendous amounts of data being generated by current sequencing projects. These efforts in analysis are being linked to new ways of visualizing computationally annotated genomes.

Animals↗

A Populus EST resource for plant functional genomics.

Trees present a life form of paramount importance for terrestrial ecosystems and human societies because of their ecological structure and physiological function and provision of energy and industrial materials. The genus Populus is the internationally accepted model for molecular tree biology. We have analyzed 102,019 Populus ESTs that clustered into 11,885 clusters and 12,759 singletons. We also provide >4,000 assembled full clone sequences to serve as a basis for the upcoming annotation of the Populus genome sequence. A public web-based EST database (POPULUSDB) provides digital expression profiles for 18 tissues that comprise the majority of differentiated organs. The coding content of Populus and Arabidopsis genomes shows very high similarity, indicating that differences between these annual and perennial angiosperm life forms result primarily from differences in gene regulation. The high similarity between Populus and Arabidopsis will allow studies of Populus to directly benefit from the detailed functional genomic information generated for Arabidopsis, enabling detailed insights into tree development and adaptation. These data will also valuable for functional genomic efforts in Arabidopsis.

Ecosystem↗

PUMA2--grid-based high-throughput analysis of genomes and metabolic pathways.

The PUMA2 system (available at http://compbio.mcs.anl.gov/puma2) is an interactive, integrated bioinformatics environment for high-throughput genetic sequence analysis and metabolic reconstructions from sequence data. PUMA2 provides a framework for comparative and evolutionary analysis of genomic data and metabolic networks in the context of taxonomic and phenotypic information. Grid infrastructure is used to perform computationally intensive tasks. PUMA2 currently contains precomputed analysis of 213 prokaryotic, 22 eukaryotic, 650 mitochondrial and 1493 viral genomes and automated metabolic reconstructions for >200 organisms. Genomic data is annotated with information integrated from >20 sequence, structural and metabolic databases and ontologies. PUMA2 supports both automated and interactive expert-driven annotation of genomes, using a variety of publicly available bioinformatics tools. It also contains a suite of unique PUMA2 tools for automated assignment of gene function, evolutionary analysis of protein families and comparative analysis of metabolic pathways. PUMA2 allows users to submit batch sequence data for automated functional analysis and construction of metabolic models. The results of these analyses are made available to the users in the PUMA2 environment for further interactive sequence analysis and annotation.

Computational Biology↗

A sequence-based identification of the genes detected by probesets on the Affymetrix U133 plus 2.0 array.

One of the biggest problems facing microarray experiments is the difficulty of translating results into other microarray formats or comparing microarray results to other biochemical methods. We believe that this is largely the result of poor gene identification. We re-identified the probesets on the Affymetrix U133 plus 2.0 GeneChip array. This identification was based on the sequence of the probes and the sequence of the human genome. Using the BLAST program, we matched probes with documented and postulated human transcripts. This resulted in the redefinition of approximately 37% of the probes on the U133 plus 2.0 array. This updated identification specifically points out where the identification is complicated by cross-hybridization from splice variants or closely related genes. More than 5000 probesets detect multiple transcripts and therefore the exact protein affected cannot be readily concluded from the performance of one probeset alone. This makes naming difficult and impacts any downstream analysis such as associating gene ontologies, mapping affected pathways or simply validating expression changes. We have now automated the sequence-based identification and can more appropriately annotate any array where the sequence on each spot is known.

Base Sequence↗

Extensive duplication and reshuffling in the Arabidopsis genome.

Systematic analysis of the Arabidopsis genome provides a basis for detailed studies of genome structure and evolution. Members of multigene families were mapped, and random sequence alignment was used to identify regions of extended similarity in the Arabidopsis genome. Detailed analysis showed that the number, order, and orientation of genes were conserved over large regions of the genome, revealing extensive duplication covering the majority of the known genomic sequence. Fine mapping analysis showed much rearrangement, resulting in a patchwork of duplicated regions that indicated deletion, insertion, tandem duplication, inversion, and reciprocal translocation. The implications of these observations for evolution of the Arabidopsis genome as well as their usefulness for analysis and annotation of the genomic sequence and in comparative genomics are discussed.

Arabidopsis↗

Gene discovery and expression profile analysis through sequencing of expressed sequence tags from different developmental stages of the chytridiomycete Blastocladiella emersonii.

Blastocladiella emersonii is an aquatic fungus of the chytridiomycete class which diverged early from the fungal lineage and is notable for the morphogenetic processes which occur during its life cycle. Its particular taxonomic position makes this fungus an interesting system to be considered when investigating phylogenetic relationships and studying the biology of lower fungi. To contribute to the understanding of the complexity of the B. emersonii genome, we present here a survey of expressed sequence tags (ESTs) from various stages of the fungal development. Nearly 20,000 cDNA clones from 10 different libraries were partially sequenced from their 5' end, yielding 16,984 high-quality ESTs. These ESTs were assembled into 4,873 putative transcripts, of which 48% presented no matches with existing sequences in public databases. As a result of Gene Ontology (GO) project annotation, 1,680 ESTs (35%) were classified into biological processes of the GO structure, with transcription and RNA processing, protein biosynthesis, and transport as prevalent processes. We also report full-length sequences, useful for construction of molecular phylogenies, and several ESTs that showed high similarity with known proteins, some of which were not previously described in fungi. Furthermore, we analyzed the expression profile (digital Northern analysis) of each transcript throughout the life cycle of the fungus using Bayesian statistics. The in silico approach was validated by Northern blot analysis with good agreement between the two methodologies.

Amino Acid Sequence↗

Repbase Update, a database of eukaryotic repetitive elements.

Repbase Update is a comprehensive database of repetitive elements from diverse eukaryotic organisms. Currently, it contains over 3600 annotated sequences representing different families and subfamilies of repeats, many of which are unreported anywhere else. Each sequence is accompanied by a short description and references to the original contributors. Repbase Update includes Repbase Reports, an electronic journal publishing newly discovered transposable elements, and the Transposon Pub, a web-based browser of selected chromosomal maps of transposable elements. Sequences from Repbase Update are used to screen and annotate repetitive elements using programs such as Censor and RepeatMasker. Repbase Update is available on the worldwide web at http://www.girinst.org/Repbase_Update.html.

Animals↗

Cloning, expression and identification of a new trehalose synthase gene from Thermobifida fusca genome.

A new open reading frame in Thermobifida fusca sequenced genome was identified to encode a new trehalose synthase, annotated as "glycosidase" in the GenBank database, by bioinformatics searching and experimental validation. The gene had a length of 1830 bp with about 65% GC content and encoded for a new trehalose synthase with 610 amino acids and deduced molecular weight of 66 kD. The high GC content seemed not to affect its good expression in E. coli BL21 in which the target protein could account for as high as 15% of the total cell proteins. The recombinant enzyme showed its optimal activities at 25 degrees and pH 6.5 when it converted substrate maltose into trehalose. However it would divert a high proportion of its substrate into glucose when the temperature was increased to 37 degrees, or when the enzyme concentration was high Its activity was not inhibited by 5 mM heavy metals such as Cu2+, Mn2+, and Zn2+ but affected by high concentration of glucose. Blasting against the database indicated that amino acid sequence of this protein had maximal 69% homology with the known trehalose synthases, and two highly conserved segments of the protein sequence were identified and their possible linkage with functions was discussed.

Actinomycetales↗

Predicting functions from protein sequences--where are the bottlenecks?

The exponential growth of sequence data does not necessarily lead to an increase in knowledge about the functions of genes and their products. Prediction of function using comparative sequence analysis is extremely powerful but, if not performed appropriately, may also lead to the creation and propagation of assignment errors. While current homology detection methods can cope with the data flow, the identification, verification and annotation of functional features need to be drastically improved.

Amino Acid Sequence↗