Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44Linked to original sources

Applied proteomics: mitochondrial proteins and effect on function.

The identification of a majority of the polypeptides in mitochondria would be invaluable because they play crucial and diverse roles in many cellular processes and diseases. The endogenous production of reactive oxygen species (ROS) is a major limiter of life as illustrated by studies in which the transgenic overexpression in invertebrates of catalytic antioxidant enzymes results in increased lifespans. Mitochondria have received considerable attention as a principal source---and target---of ROS. Mitochondrial oxidative stress has been implicated in heart disease including myocardial preconditioning, ischemia/reperfusion, and other pathologies. In addition, oxidative stress in the mitochondria is associated with the pathogenesis of Alzheimer's disease, Parkinson's disease, prion diseases, and amyotrophic lateral sclerosis (ALS) as well as aging itself. The rapidly emerging field of proteomics can provide powerful strategies for the characterization of mitochondrial proteins. Current approaches to mitochondrial proteomics include the creation of detailed catalogues of the protein components in a single sample or the identification of differentially expressed proteins in diseased or physiologically altered samples versus a reference control. It is clear that for any proteomics approach prefractionation of complex protein mixtures is essential to facilitate the identification of low-abundance proteins because the dynamic range of protein abundance within cells has been estimated to be as high as 10(7). The opportunities for identification of proteins directly involved in diseases associated with or caused by mitochondrial dysfunction are compelling. Future efforts will focus on linking genomic array information to actual protein levels in mitochondria.

Animals↗

New insights into the rat spermatogonial proteome: identification of 156 additional proteins.

Despite the essential role played by spermatogonia in testicular function, little is known about these cells. To improve our understanding of their biology, our group recently identified a set of 53 spermatogonial proteins using two-dimensional (2-D) gel electrophoresis and mass spectrometry. To continue this work, we investigated a subset of the spermatogonial proteome using narrow range immobilized pH gradients to favor the detection of less abundant proteins. A 2-D reference map of spermatogonia in the pH range 4-9 was created, and protein entities fractionated in a pH 5-6 2-D gel were further processed for protein identification. A new set of 156 polypeptides was identified by peptide mass fingerprinting and tandem mass spectrometry. These polypeptides corresponded to 102 different proteins, which reflect the complexity of post-translational modifications. Seventy-nine of these proteins were identified for the first time in spermatogonia. All identified proteins were classified into functional groups. This work represents a first step toward the establishment of a systematic spermatogonia protein database.

Animals↗

Comparison of knowledge-based and distance geometry approaches for generation of molecular conformations.

A knowledge-based approach for generating conformations of molecules has been developed. The method described here provides a good sampling of the molecule's conformational space by restricting the generated conformations to those consistent with the reference database. The present approach, internally named et for enumerate torsions, differs from previous database-mining approaches by employing a library of much larger substructures while treating open chains, rings, and combinations of chains and rings in the same manner. In addition to knowledge in the form of observed torsion angles, some knowledge from the medicinal chemist is captured in the form of which substructures are identified. The knowledge-based approach is compared to Blaney et al.'s distance geometry (DG) algorithm for sampling the conformational space of molecules. The structures of 113 protein-bound molecules, determined by X-ray crystallography, were used to compare the methods. The present knowledge-based approach (i) generates conformations closer to the experimentally determined conformation, (ii) generates them sooner, and (iii) is significantly faster than the DG method.

Algorithms↗

Homologous sequences in steroidogenic enzymes, steroid receptors and a steroid binding protein suggest a consensus steroid-binding sequence.

The amino acid sequences of two steroidogenic enzymes, P450c17 (steroid 17 alpha-hydroxylase/17,20 lyse) and P450c21 (steroid 21-hydroxylase), are only 28.9% identical. However, these proteins share a region of 21 amino acids bearing 17 identical residues, which we previously suggested may represent the steroid binding site. We assembled a sequence database of known steroid-binding proteins and searched this with the sequence of this 21 amino acid region. The steroidogenic enzymes, P450c17, P450c21, P450scc (the cholesterol side-chain cleavage enzyme), and P450c11 (steroid 11 beta/18-hydroxylase) share a subregion of 17 amino acids having at least 15 identical residues. Related sequences were identified in a computerized search of the available sequences of steroid hormone receptors and binding proteins. These sequences were invariably found within larger domains previously associated with steroid binding. From these we propose a more general consensus sequence of LPLLL +/- 000KDRE0LKRL +/- PV, where +/- refers to any charged amino acid, and 0 refers to an uncharged amino acid. This consensus sequence predicts 147 or 187 total amino acids in 11 human proteins examined (78.6%). An equivalent degree of sequence identity, 178 of 221 amino acids (80.5%) was found among 13 animal homologs of these human proteins. The ability of this consensus sequence to predict 325 of 408 amino acids (79.7%) strongly suggests this sequence is necessary, if not sufficient, for a steroid binding site in many proteins. Lecithin-cholesterol acetyl transferase, cholesterol ester transfer protein, and steroid sulfatase did not have sequences similar to our consensus sequence.

Amino Acid Sequence↗

The chromosome-level genome assembly and annotation of the silver-lipped pearl oyster, Pinctada maxima.

The silver-lipped pearl oyster (Pinctada maxima) is a valuable tropical aquaculture species, playing a crucial economic role in the global pearl industry. However, the lack of genomic reference limits our in-depth understanding of this species in genome-based breeding, conservation, evolution and adaptation. Here, annotated chromosome-level reference genome for P. maxima was generated by integrating PacBio long-read sequencing, Illumina short-read sequencing, and Hi-C sequencing data. The total genome size is 1,264.93&#x2009;Mb, with contig N50 and scaffold N50 of 649&#x2009;kb and 89.19&#x2009;Mb, respectively. The majority (97.94%) of the assembled genome was anchored to the 14 chromosomes by Hi-C analysis. The relatively high genome completeness was observed, with 97.38% (metazoa_odb10 database) and 95.26% (mollusca_odb10 database) in BUSCO analysis. Genome annotation revealed approximately 65.46% of the repeat sequences and 26,315 protein-coding genes. Comparative genome analysis revealed 28 expanded and 48 contracted families (p&#x2009;<&#x2009;0.05) in P. maxima, with 3.2% of genes (894) being species-specific. This chromosome-level genome serves as an essential resource for research in evolutionary genomics, phylogenetics, and biomineralization.

Animals↗

Optimizing the benefits and minimizing the risks of enteral nutrition in the critically ill: role of small bowel feeding.

BACKGROUND: Strategies that maximize the delivery of enteral nutrition while minimizing the associated risks have the potential to improve the outcomes of critically ill patients. By delivering enteral feeds in the small bowel, beyond the pylorus, the frequency of regurgitation and the risk of aspiration is thought to be decreased while at the same time, nutrient delivery is maximized. The purpose of this paper is to systematically review those studies that compare gastric with small bowel feeding. METHODS: We searched computerized bibliographic databases, personal files, and relevant reference lists to identify eligible studies. Only randomized, clinical trials of critically ill patients that compared small bowel and gastric feedings were included in this review. In an independent fashion, relevant data on the methodology and outcomes of primary studies were abstracted in duplicate. RESULTS: There were 10 studies that met the inclusion criteria for this review. In 1 study, small bowel feeding was associated with a reduction in gastroesophageal regurgitation and a trend toward reduced pulmonary aspiration. Several studies document that small bowel feeding was associated with an increase in protein and calories delivered and a shorter time to target dose of nutrition. Compared with gastric feeding, when the results of 7 randomized trials were aggregated statistically, small bowel feeding was associated with a reduction in pneumonia (relative risk, 0.76; 95% confidence intervals, 0.59, 0.99). There was no difference in mortality rates between the 2 groups. CONCLUSIONS: Small bowel feeding may be associated with a reduction in gastroesophageal regurgitation, an increase in nutrient delivery, a shorter time to achieve desired target nutrition, and a lower rate of ventilator-associated pneumonia.

Critical Illness↗

FLEXS: a method for fast flexible ligand superposition.

If no structural information about a particular target protein is available, methods of rational drug design try to superimpose putative ligands with a given reference, e.g., an endogenous ligand. The goal of such structural alignments is, on the one hand, to approximate the binding geometry and, on the other hand, to provide a relative ranking of the ligands with respect to their similarity. An accurate superposition is the prerequisite of subsequent exploitation of ligand data by either 3D QSAR analyses, pharmacophore hypotheses, or receptor modeling. We present the automatic method FLEXS for structurally superimposing pairs of ligands, approximating their putative binding site geometry. One of the ligands is treated as flexible, while the other one, used as a reference, is kept rigid. FLEXS is an incremental construction procedure. The molecules to be superimposed are partitioned into fragments. Starting with placements of a selected anchor fragment, computed by two alternative approaches, the remaining fragments are added iteratively. At each step, flexibility is considered by allowing the respective added fragment to adopt a discrete set of conformations. The mean computing time per test case is about 1:30 min on a common-day workstation. FLEXS is fast enough to be used as a tool for virtual ligand screening. A database of typical drug molecules has been screened for potential fibrinogen receptor antagonists. FLEXS is capable of retrieving all ligands assigned to platelet aggregation properties among the first 20 hits. Furthermore, the program suggests additional interesting candidates, likely to be active at the same receptor. FLEXS proves to be superior to commonly used retrieval techniques based on 2D fingerprint similarities. The accuracy of computed superpositions determines the relevance of subsequently performed ligand analyses. In order to validate the quality of FLEXS alignments, we attempted to reproduce a set of 284 mutual superpositions derived from experimental data on 76 protein-ligand complexes of 14 proteins. The ligands considered cover the whole range of drug-size molecules from 18 to 158 atoms (PDB codes: 3ptb, 2er7). The performance of the algorithm critically depends on the sizes of the molecules to be superimposed. The limitations are clearly demonstrated with large peptidic inhibitors in the HIV and the endothiapepsin data set. Problems also occur in the presence of multiple binding modes (e.g., elastase and human rhinovirus). The most convincing results are achieved with small- and medium-sized molecules (as, e.g., the ligands of trypsin, thrombin, and dihydrofolate reductase). In more than half of the entire test set, we achieve rms deviations between computed and observed alignment of below 1.5 A. This underlines the reliability of FLEXS-generated alignments.

Algorithms↗

An evidence-based approach to earlier initiation of dialysis.

The objective was to review evidence addressing the optimal time to initiate dialysis treatment. The database was derived from an evidence-based review of the medical literature and from the Canada-United States peritoneal dialysis study. The publications were divided into (1) those addressing the clinical impact of early versus late referral to a dialysis program; (2) those evaluating the association between residual renal function at initiation of dialysis and the concurrent nutritional status; (3) those evaluating the association between residual renal function at initiation of dialysis and subsequent clinical outcomes, including patient survival. There were five studies evaluating early versus late referral, three cohort design and two case-control design. Late referrals had worse outcomes than early referrals. The former had more serious comorbidity and many had been noncompliant with follow-up. The latter were more likely to have hereditary renal disease. Renal function was slightly worse at initiation among those referred late. Three studies addressed the association between renal function at initiation of dialysis and concurrent nutritional status. Two showed decreased protein intake with diminished glomerular filtration rate (GFR). Poor nutritional status is associated with decreased patient survival among both incident and prevalent dialysis patients. The third study reported excellent patient survival among patients with late initiation of dialysis. These patients had received a supplemented low-protein diet and were not malnourished at initiation of dialysis. Three groups have studied the association between GFR at initiation of dialysis and clinical outcomes. Decreased GFR at initiation of dialysis is associated with a increased probability of hospitalization and death. None of these studies has used the rigorous randomized clinical trial design, and they are therefore subject to bias. Referral time bias, comorbidity, patient compliance, and starting time bias are potential confounders. A randomized clinical trial is required to resolve this important issue. However, there is sufficient evidence to justify initiation of dialysis at a Ccr of 9 to 14 mL/min if there is any clinical or laboratory evidence of malnutrition.

Bias↗

Mutations in the SLCO1B3 gene affecting the substrate specificity of the hepatocellular uptake transporter OATP1B3 (OATP8).

OBJECTIVE: Hepatocellular uptake transporters are involved in the hepatobiliary elimination of endogenous and xenobiotic substances. Mutations in genes encoding these uptake transporters may be key determinants of interindividual variability in hepatobiliary elimination and drug disposition. Our aim was to investigate the functional consequences of mutations in the SLCO1B3 gene encoding the hepatic uptake transporter for organic anions OATP1B3, formerly termed OATP8. METHODS: Mutations occurring in Caucasian Europeans and observed in databases were introduced into the SLCO1B3 cDNA and the consequences were analyzed in stably transfected canine MDCKII cells and human HEK293 cells. The functional consequences were examined for two frequent polymorphisms SLCO1B3-334T>G, encoding OATP1B3-S112A (allelic frequency of 74%) and SLCO1B3-699G>A, encoding OATP1B3-M233I (allelic frequency of 71%) and one rare polymorphism SLCO1B3-1564G>T, encoding OATP1B3-G522C (allelic frequency of 1.9%) and one artificial mutation SLCO1B3-1748G>A, encoding OATP1B3-G583E. RESULTS: OATP1B3-S112A, OATP1B3-M233I, and the OATP1B3 protein corresponding to the reference sequence (accession NM_019844), showed a comparable lateral localization in stably transfected MDCKII cells, whereas OATP1B3-G522C and OATP1B3-G583E proteins were retained intracellularly. Both latter amino acid substitutions abolished the transport of bile acids mediated by OATP1B3, whereas other substrates, like bromosulfophthalein, were transported by all polymorphic variants of the protein. CONCLUSIONS: The functional consequences of three polymorphisms and one artificial mutation include differences in the localization and in transport characteristics of several OATP1B3 proteins. This study demonstrates the importance of the analysis of genetic variations in genes encoding transport proteins for the understanding of individual variations in the hepatobiliary elimination of substances.

Amino Acid Sequence↗

Identification and preliminary characterization of a 75-kDa hemin- and hemoglobin-binding outer membrane protein of Actinobacillus pleuropneumoniae serotype 1.

The reference strains representing serotypes 1 to 12 of Actinobacillus pleuropneumoniae biotype 1 were examined for their ability to utilize porcine hemoglobin (Hb) or porcine hemin (Hm) as iron sources for growth. In a growth promotion assay, all of the reference strains were able to use porcine Hb, and all strains except 2 were able to use porcine Hm. Using a preliminary characterization procedure with Hm- or Hb-agarose, Hm- and Hb-binding outer membrane proteins (OMPs) of approximately 75 kDa were isolated from A. pleuropneumoniae serotype 1 strain 4074 grown under iron-restricted conditions. Matrix-assisted laser desorption ionization/time-of-flight (MALDI-TOF) analysis revealed a number of common tryptic peptides between the Hb-agarose- and Hm-agarose-purified 75 kDa OMPs, strongly suggesting that these peptides originate from the same protein. A database search of these peptide sequences revealed identities with proteins from various Gram-negative bacteria, including iron-regulated OMPs, transporter proteins, as well as TonB-dependent receptors. Taken together, our data suggest that A. pleuropneumoniae synthesizes potential Hm- and Hb-binding proteins that could be implicated in the iron uptake from porcine Hb and Hm.

Actinobacillus pleuropneumoniae↗

Characterization of a novel gene, STAG1/PMEPA1, upregulated in renal cell carcinoma and other solid tumors.

Using differential display-polymerase chain reaction, we identified a novel gene sequence, designated solid tumor-associated gene 1 (STAG1), that is upregulated in renal cell carcinoma (RCC). The full-length cDNA (4839 bp) encompassed the recently reported androgen-regulated prostatic cDNA PMEPA1, and so we refer to this gene as STAG1/PMEPA1. Two STAG1/PMEPA1 mRNA transcripts of approximately 2.7 and 5 kb, with identical coding regions but variant 3' untranslated regions, were predominantly expressed in normal prostate tissue and at lower levels in the ovary. The expression of this gene was upregulated in 87% of RCC samples and also was upregulated in stomach and rectal adenocarcinomas. In contrast, STAG1/PMEPA1 expression was barely detectable in leukemia and lymphoma samples. Analysis of expressed sequence tag databases showed that STAG1/PMEPA1 also was expressed in pancreatic, endometrial, and prostatic adenocarcinomas. The STAG1/PMEPA1 cDNA encodes a 287-amino-acid protein containing a putative transmembrane domain and motifs that suggest that it may bind src homology 3- and tryptophan tryptophan domain-containing proteins. This protein shows 67% identity to the protein encoded by the chromosome 18 open reading frame 1 gene. Translation of STAG1/PMEPA1 mRNA in vitro showed two products of 36 and 39 kDa, respectively, suggesting that translation may initiate at more than one site. Comparison to genomic clones showed that STAG1/PMEPA1 was located on chromosome 20q13 between microsatellite markers D20S183 and D20S173 and spanned four exons and three introns. The upregulation of this gene in several solid tumors indicated that it may play an important role in tumorigenesis.

Amino Acid Sequence↗

Trypsin-based monolithic bioreactor coupled on-line with LC/MS/MS system for protein digestion and variant identification in standard solutions and serum samples.

The applicability of a trypsin-based monolithic bioreactor coupled on-line with LC/MS/MS for rapid proteolytic digestion and protein identification is here described. Dilute samples are passed through the bioreactor for generation of proteolytic fragments in less than 10 min. After digestion and peptide separation, electrospray ionization tandem mass spectrometry is used to generate a peptide map and to identify proteolytic peptides by correlating their fragmentation spectra with amino acid sequences from a protein database. By digesting picomoles of proteins sufficient data from ESI and MS/MS were obtained to unambiguously identify proteins alone and in serum samples. This approach was also extended to locate mutation sites in beta-lactoglobulin A and B variants.

Amino Acid Sequence↗

Functional genomics and enzyme evolution. Homologous and analogous enzymes encoded in microbial genomes.

Computational analysis of complete genomes, followed by experimental testing of emerging hypotheses--the area of research often referred to as 'functional genomics'--aims at deciphering the wealth of information contained in genome sequences and at using it to improve our understanding of the mechanisms of cell function. This review centers on the recent progress in the genome analysis with special emphasis on the new insights in enzyme evolution. Standard methods of predicting functions for new proteins are listed and the common errors in their application are discussed. A new method of improving the functional predictions is introduced, based on a phylogenetic approach to functional prediction, as implemented in the recently constructed Clusters of Orthologous Groups (COG) database (available at http:@www.ncbi.nlm.nih.gov/COG). This approach provides a convenient way to characterize the protein families (and metabolic pathways) that are present or absent in any given organism. Comparative analysis of microbial genomes based on this approach shows that metabolic diversity generally correlates with the genome size-parasitic bacteria code for fewer enzymes and lesser number of metabolic pathways than their free-living relatives. Comparison of different genomes reveals another evolutionary trend, the non-orthologous gene displacement of some enzymes by unrelated proteins with the same cellular function. An examination of the phylogenetic distribution of such cases provides new clues to the problems of biochemical evolution, including evolution of glycolysis and the TCA cycle.

Databases, Factual↗

Report of the International Equine Gene Mapping Workshop: male linkage map.

The goal of the First International Equine Gene Mapping Workshop, held in 1995, was the construction of a low density, male linkage map for the horse. For this purpose, the International Horse Reference Family Panel (IHRFP) was established, consisting of 12 paternal half-sib families with 448 half-sib offspring provided by 10 laboratories. Blood samples were collected and DNA extracted in each laboratory and sent to the Lexington laboratory (KY, USA) for dispatch in aliquots to 14 typing laboratories. In total, 161 markers (144 microsatellites, seven blood groups and 10 proteins) were tested for all families for which the sire was heterozygous. Genealogies and typing data were sent for analysis to the INRA laboratory (Jouy-en-Josas, France) according to a specific format and entered into a database with input verification and output processes. Linkage analysis was performed with the CRIMAP program. Significant linkage was detected for 124 loci, of which 95 were unambiguously ordered using a multipoint analysis with an average spacing of 14.2 CM. These loci were distributed among 29 linkage groups. A more comprehensive analysis including synteny group data and FISH data suggested that 26 autosomes out of 31 are covered. The complete map spans 936 CM.

Animals↗

A hybrid approach for addressing ring flexibility in 3D database searching.

A hybrid approach for flexible 3D database searching is presented that addresses the problem of ring flexibility. It combines the explicit storage of up to 25 multiple conformations of rings, with up to eight atoms, generated by the 3D structure generator CORINA with the power of a torsional fitting technique implemented in the 3D database system UNITY. A comparison with the original UNITY approach, using a database with about 130,000 entries and five different pharmacophore queries, was performed. The hybrid approach scored, on an average, 10-20% more hits than the reference run. Moreover, specific problems with unrealistic hit geometries produced by the original approach can be excluded. In addition, the influence of the maximum number of ring conformations per molecule was investigated. An optimal number of 10 conformations per molecule is recommended.

Anti-Arrhythmia Agents↗

[Transcriptomes for serial analysis of gene expression].

The availability of the sequences for whole genomes is changing our understanding of cell biology. Functional genomics refers to the comprehensive analysis, at the protein level (proteome) and at the mRNA level (transcriptome) of all events associated with the expression of whole sets of genes. New methods have been developed for transcriptome analysis. Serial Analysis of Gene Expression (SAGE) is based on the massive sequential analysis of short cDNA sequence tags. Each tag is derived from a defined position within a transcript. Its size (14 bp) is sufficient to identify the corresponding gene and the number of times each tag is observed provides an accurate measurement of its expression level. Since tag populations can be widely amplified without altering their relative proportions, SAGE may be performed with minute amounts of biological extract. Dealing with the mass of data generated by SAGE necessitates computer analysis. A software is required to automatically detect and count tags from sequence files. Criterias allowing to assess the quality of experimental data can be included at this stage. To identify the corresponding genes, a database is created registering all virtual tags susceptible to be observed, based on the present status of the genome knowledge. By using currently available database functions, it is easy to match experimental and virtual tags, thus generating a new database registering identified tags, together with their expression levels. As an open system, SAGE is able to reveal new, yet unknown, transcripts. Their identification will become increasingly easier with the progress of genome annotation. However, their direct characterization can be attempted, since tag information may be sufficient to design primers allowing to extend unknown sequences. A major advantage of SAGE is that, by measuring expression levels without reference to an arbitrary standard, data are definitively acquired and cumulative. All publicly available data can thus be stored in a unique database, facilitating whole-genome analysis of differential expression between cell types, normal and diseased samples, or samples with and without drug treatment. SAGE data are readily amenable to statistical comparisons, allowing to determine the level of confidence of the observed variations. A major limitation of SAGE is that, because each analysis is obligatory performed on the whole set of expressed genes, it can hardly be performed on multiple samples, for example in kinetics studies or to compare the effects of large numbers of drugs. To overcome this limitation, high-throughput detection of a subset of mRNAs is more rapidly performed by parallel hybridization of mRNAs on arrays of nucleic acids immobilized on solid supports. From this point of view, a SAGE platform is a powerful instrument for selecting the most informative subset of genes, assembling them to design microarrays dedicated to a specific problem and calibrating measurement by comparison with a standard cell model for which SAGE data are available. This approach is an attractive alternative to strategies based exclusively on pangenomic arrays. A very large amount of SAGE data are already available and the problem is now to extract their biological meaning. Knowledge on metabolic pathways is already organized so that its successful integration in a SAGE platform can be undertaken. For other cell components and pathways, the problem lies on the lack of controlled vocabulary to describe gene activities, starting form a clear definition of the concept of biological function itself. Progress in gene and cell ontology is expected to facilitate computer-based extraction of biological knowledge from existing and forthcoming SAGE data.

Animals↗

A comparison of scoring functions for protein sequence profile alignment.

MOTIVATION: In recent years, several methods have been proposed for aligning two protein sequence profiles, with reported improvements in alignment accuracy and homolog discrimination versus sequence-sequence methods (e.g. BLAST) and profile-sequence methods (e.g. PSI-BLAST). Profile-profile alignment is also the iterated step in progressive multiple sequence alignment algorithms such as CLUSTALW. However, little is known about the relative performance of different profile-profile scoring functions. In this work, we evaluate the alignment accuracy of 23 different profile-profile scoring functions by comparing alignments of 488 pairs of sequences with identity < or =30% against structural alignments. We optimize parameters for all scoring functions on the same training set and use profiles of alignments from both PSI-BLAST and SAM-T99. Structural alignments are constructed from a consensus between the FSSP database and CE structural aligner. We compare the results with sequence-sequence and sequence-profile methods, including BLAST and PSI-BLAST. RESULTS: We find that profile-profile alignment gives an average improvement over our test set of typically 2-3% over profile-sequence alignment and approximately 40% over sequence-sequence alignment. No statistically significant difference is seen in the relative performance of most of the scoring functions tested. Significantly better results are obtained with profiles constructed from SAM-T99 alignments than from PSI-BLAST alignments. AVAILABILITY: Source code, reference alignments and more detailed results are freely available at http://phylogenomics.berkeley.edu/profilealignment/

Algorithms↗

MIPS Arabidopsis thaliana Database (MAtDB): an integrated biological knowledge resource for plant genomics.

Arabidopsis thaliana is the most widely studied model plant. Functional genomics is intensively underway in many laboratories worldwide. Beyond the basic annotation of the primary sequence data, the annotated genetic elements of Arabidopsis must be linked to diverse biological data and higher order information such as metabolic or regulatory pathways. The MIPS Arabidopsis thaliana database MAtDB aims to provide a comprehensive resource for Arabidopsis as a genome model that serves as a primary reference for research in plants and is suitable for transfer of knowledge to other plants, especially crops. The genome sequence as a common backbone serves as a scaffold for the integration of data, while, in a complementary effort, these data are enhanced through the application of state-of-the-art bioinformatics tools. This information is visualized on a genome-wide and a gene-by-gene basis with access both for web users and applications. This report updates the information given in a previous report and provides an outlook on further developments. The MAtDB web interface can be accessed at http://mips.gsf.de/proj/thal/db.

Arabidopsis↗