Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Databases, Genetic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Database and analyses of known alternatively spliced genes in plants.

Alternative splicing is an important cellular mechanism that increases the diversity of gene products. The number of alternatively spliced genes reported so far in plants is much smaller than that in mammals, but is increasing as a result of the explosive growth of available EST and genomic sequences. We have searched for all alternatively spliced genes reported in GenBank and PubMed in all plant species under Viridiplantae. After careful merging and manual review of the search results, we obtained a comprehensive, high-quality collection of 168 genes reported to be alternatively spliced in plants, spanning 44 plant species (March 22, 2003 update). We developed a relational database with Web-based user interface to store and present the data, named the Plant Alternative Splicing Database (PASDB), freely available at http://pasdb.genomics.org.cn. We analyzed the functional categories that these genes belong to using the Gene Ontology. We also analyzed in detail the biological roles and gene structures of the four genes that are known to be alternatively spliced in more than one plant species. Finally, we studied the structural features of the splice sites in the alternatively spliced genes.

Alternative Splicing↗

Systematic analysis of human kinase genes: a large number of genes and alternative splicing events result in functional and structural diversity.

BACKGROUND: Protein kinases are a well defined family of proteins, characterized by the presence of a common kinase catalytic domain and playing a significant role in many important cellular processes, such as proliferation, maintenance of cell shape, apoptosis. In many members of the family, additional non-kinase domains contribute further specialization, resulting in subcellular localization, protein binding and regulation of activity, among others. About 500 genes encode members of the kinase family in the human genome, and although many of them represent well known genes, a larger number of genes code for proteins of more recent identification, or for unknown proteins identified as kinase only after computational studies. RESULTS: A systematic in silico study performed on the human genome, led to the identification of 5 genes, on chromosome 1, 11, 13, 15 and 16 respectively, and 1 pseudogene on chromosome X; some of these genes are reported as kinases from NCBI but are absent in other databases, such as KinBase. Comparative analysis of 483 gene regions and subsequent computational analysis, aimed at identifying unannotated exons, indicates that a large number of kinase may code for alternately spliced forms or be incorrectly annotated. An InterProScan automated analysis was performed to study domain distribution and combination in the various families. At the same time, other structural features were also added to the annotation process, including the putative presence of transmembrane alpha helices, and the cystein propensity to participate into a disulfide bridge. CONCLUSION: The predicted human kinome was extended by identifying both additional genes and potential splice variants, resulting in a varied panorama where functionality may be searched at the gene and protein level. Structural analysis of kinase proteins domains as defined in multiple sources together with transmembrane alpha helices and signal peptide prediction provides hints to function assignment. The results of the human kinome analysis are collected in the KinWeb database, available for browsing and searching over the internet, where all results from the comparative analysis and the gene structure annotation are made available, alongside the domain information. Kinases may be searched by domain combinations and the relative genes may be viewed in a graphic browser at various level of magnification up to gene organization on the full chromosome set.

Algorithms↗

Effect of dataset selection on the topological interpretation of protein interaction networks.

BACKGROUND: Studies of the yeast protein interaction network have revealed distinct correlations between the connectivity of individual proteins within the network and the average connectivity of their neighbours. Although a number of biological mechanisms have been proposed to account for these findings, the significance and influence of the specific datasets included in these studies has not been appreciated adequately. RESULTS: We show how the use of different interaction data sets, such as those resulting from high-throughput or small-scale studies, and different modelling methodologies for the derivation pair-wise protein interactions, can dramatically change the topology of these networks. Furthermore, we show that some of the previously reported features identified in these networks may simply be the result of experimental or methodological errors and biases. CONCLUSION: When performing network-based studies, it is essential to define what is meant by the term "interaction" and this must be taken into account when interpreting the topologies of the networks generated. Consideration must be given to the type of data included and appropriate controls that take into account the idiosyncrasies of the data must be selected.

Computational Biology↗

Diversity of picoplanktonic prasinophytes assessed by direct nuclear SSU rDNA sequencing of environmental samples and novel isolates retrieved from oceanic and coastal marine ecosystems.

Picoplanktonic prasinophytes are well represented in culture collections and marine samples. In order to better characterize this ecologically important group, we compared the phylogenetic diversity of picoplanktonic prasinophyte strains available at the Roscoff Culture Collection (RCC) and that of nuclear SSU rDNA sequences from environmental clone libraries obtained from oceanic and coastal ecosystems. Among the 570 strains avalaible, 91 belonged to prasinophytes, 65 were partially sequenced, and we obtained the entire SSU rDNA sequence for a selection of 14 strains. Within the 18 available environmental clone libraries, the prasinophytes accounted for 12% of the total number of clones retrieved (142 partial sequences in total), and we selected 9 clones to obtain entire SSU rDNA sequence. Using this approach, we obtained a subsequent genetic database that revealed the presence of seven independent lineages among prasinophytes, including a novel clade (clade VII). This new clade groups the genus Picocystis, two unidentified coccoid strains, and 4 environmental sequences. For each of these seven lineages, at least one representative is available in culture. The three picoplanktonic genera Ostreococcus, Micromonas, and Bathycoccus (order Mamiellales), were the best represented prasinophytes both in cultures and genetic libraries. SSU rDNA phylogenetic analyses suggest that the genus Bathycoccus forms a very homogeneous group. In contrast, the genera Micromonas and Ostreococcus turned out to be quite complex, consisting of three and four independent lineages, respectively. This report of the overall diversity of picoeukaryotic prasinophytes reveals a group of ecologically important and diverse marine microorganims that are well represented by isolated cultures.

Animals↗

Selecting protein targets for structural genomics of Pyrobaculum aerophilum: validating automated fold assignment methods by using binary hypothesis testing.

Three-dimensional protein folds were assigned to all ORFs of the recently sequenced genome of the hyperthermophilic archaeon Pyrobaculum aerophilum. Binary hypothesis testing was used to estimate a confidence level for each assignment. A separate test was conducted to assign a probability for whether each sequence has a novel fold-i.e., one that is not yet represented in the experimental database of known structures. Of the 2,130 predicted nontransmembrane proteins in this organism, 916 matched a fold at a cumulative 90% confidence level, and 245 could be assigned at a 99% confidence level. Likewise, 286 proteins were predicted to have a previously unobserved fold with a 90% confidence level, and 14 at a 99% confidence level. These statistically based tools are combined with homology searches against the Online Mendelian Inheritance in Man (OMIM) human genetics database and other protein databases for the selection of attractive targets for crystallographic or NMR structure determination. Results of these studies have been collated and placed at http://www.doe-mbi.ucla.edu/people/parag/P A_HOME/, the University of California, Los Angeles-Department of Energy Pyrobaculum aerophilum web site.

Algorithms↗

Unusual polymorphisms in human immunodeficiency virus type 1 associated with nonprogressive infection.

Factors accounting for long-term nonprogression may include infection with an attenuated strain of human immunodeficiency virus type 1 (HIV-1), genetic polymorphisms in the host, and virus-specific immune responses. In this study, we examined eight individuals with nonprogressing or slowly progressing HIV-1 infection, none of whom were homozygous for host-specific polymorphisms (CCR5-Delta32, CCR2-64I, and SDF-1-3'A) which have been associated with slower disease progression. HIV-1 was recovered from seven of the eight, and recovered virus was used for sequencing the full-length HIV-1 genome; full-length HIV-1 genome sequences from the eighth were determined following amplification of viral sequences directly from peripheral blood mononuclear cells (PBMC). Longitudinal studies of one individual with HIV-1 that consistently exhibited a slow/low growth phenotype revealed a single amino acid deletion in a conserved region of the gp41 transmembrane protein that was not seen in any of 131 envelope sequences in the Los Alamos HIV-1 sequence database. Genetic analysis also revealed that five of the eight individuals harbored HIV-1 with unusual 1- or 2-amino-acid deletions in the Gag sequence compared to subgroup B Gag consensus sequences. These deletions in Gag have either never been observed previously or are extremely rare in the database. Three individuals had deletions in Nef, and one had a 4-amino-acid insertion in Vpu. The unusual polymorphisms in Gag, Env, and Nef described here were also found in stored PBMC samples taken 3 to 11 years prior to, or in one case 4 years subsequent to, the time of sampling for the original sequencing. In all, seven of the eight individuals exhibited one or more unusual polymorphisms; a total of 13 unusual polymorphisms were documented in these seven individuals. These polymorphisms may have been present from the time of initial infection or may have appeared in response to immune surveillance or other selective pressures. Our results indicate that unusual, difficult-to-revert polymorphisms in HIV-1 can be found associated with slow progression or nonprogression in a majority of such cases.

Amino Acid Sequence↗

Exploring the genome of Trypanosoma vivax through GSS and in silico comparative analysis.

A survey of the Trypanosoma vivax genome was carried out by the genome sequence survey (GSS) approach resulting in 1,086 genomic sequences. A total of 455 high-quality GSS sequences were generated, consisting of 331 non-redundant sequences distributed in 264 singlets and 67 clusters in a total of 135.5 Kb of the T. vivax genome. The estimation of the overall G+C content, and the prediction of the presence of ORFs and putative genes were carried out using the Glimmer and Jemboss packages. Analysis of the obtained sequences was carried out by BLAST programs against 12 different databases and also using the Conserved Domain Database, InterProScan, and tRNAscan-SE. Along with the existing 23 T. vivax entries in the GenBank, the 32 putative genes predicted and the 331 non-redundant GSS sequences reported herein represent new potential markers for the development of PCRbased assays for specific diagnosis and typing of Trypanosoma vivax.

Algorithms↗

Enhanced position weight matrices using mixture models.

MOTIVATION: Positional weight matrix (PWM) is derived from a set of experimentally determined binding sites. Here we explore whether there exist subclasses of binding sites and if the mixture of these subclass-PWMs can improve the binding site prediction. Intuitively, the subclasses correspond to either distinct binding preference of the same transcription factor in different contexts or distinct subtypes of the transcription factor. AVAILABILITY: We report an Expectation Maximization algorithm adapting the mixture model of Baily and Elkan. We assessed the relative merit of using two subclass-PWMs. The resulting PWMs were evaluated with respect to preferred conservation (relative to mouse) of potential sites in human promoters and expression coherence of the potential target genes. Based on 64 JASPAR vertebrate PWMs, 61-81% of the cases resulted in a higher conservation using the mixture model. Also in 98% of the cases the expression coherence was higher for the target genes of one of the subclass-PWMs. Our analysis of Reb1 sites is consistent with previously discovered subtypes using independent methods. Additionally application of our method to mutated sites for transcription factor LEU3 reveals subclasses that segregate into strongly binding and weakly binding sites with P-value of 0.008. This is the first study which attempts to quantify the subtly different binding specificities of a transcription factor on a large scale and suggests the use of a mixture of PWMs, instead of the current practice of using a single PWM, for a transcription factor.

Algorithms↗

Calibrating E-values for hidden Markov models using reverse-sequence null models.

MOTIVATION: Hidden Markov models (HMMs) calculate the probability that a sequence was generated by a given model. Log-odds scoring provides a context for evaluating this probability, by considering it in relation to a null hypothesis. We have found that using a reverse-sequence null model effectively removes biases owing to sequence length and composition and reduces the number of false positives in a database search. Any scoring system is an arbitrary measure of the quality of database matches. Significance estimates of scores are essential, because they eliminate model- and method-dependent scaling factors, and because they quantify the importance of each match. Accurate computation of the significance of reverse-sequence null model scores presents a problem, because the scores do not fit the extreme-value (Gumbel) distribution commonly used to estimate HMM scores' significance. RESULTS: To get a better estimate of the significance of reverse-sequence null model scores, we derive a theoretical distribution based on the assumption of a Gumbel distribution for raw HMM scores and compare estimates based on this and other distribution families. We derive estimation methods for the parameters of the distributions based on maximum likelihood and on moment matching (least-squares fit for Student's t-distribution). We evaluate the modeled distributions of scores, based on how well they fit the tail of the observed distribution for data not used in the fitting and on the effects of the improved E-values on our HMM-based fold-recognition methods. The theoretical distribution provides some improvement in fitting the tail and in providing fewer false positives in the fold-recognition test. An ad hoc distribution based on assuming a stretched exponential tail does an even better job. The use of Student's t to model the distribution fits well in the middle of the distribution, but provides too heavy a tail. The moment-matching methods fit the tails better than maximum-likelihood methods. AVAILABILITY: Information on obtaining the SAM program suite (free for academic use), as well as a server interface, is available at http://www.soe.ucsc.edu/research/compbio/sam.html and the open-source random sequence generator with varying compositional biases is available at http://www.soe.ucsc.edu/research/compbio/gen_sequence

Algorithms↗

ATID: a web-oriented database for collection of publicly available alternative translational initiation events.

SUMMARY: Alternative translational initiation is an important cellular mechanism contributing to the diversity of protein products and functions. We develop a database that provides a comprehensive collection of alternative translational initiation events. The purpose of this alternative translational initiation database (ATID) is to facilitate the systematic study of alternative translational initiation of genes. The current version of database contains 300 genes from Homo sapiens, Mus musculus and other species. Each of the genes has two or more isoforms due to alternative translational initiation. Resources in ATID, including gene information, alternative products of genes and domain structures of isoforms, are provided through a user-friendly web interface. AVAILABILITY: The ATID database is available for public use at http://bioinfo.au.tsinghua.edu.cn/atie/.

Amino Acid Sequence↗

BioThesaurus: a web-based thesaurus of protein and gene names.

UNLABELLED: BioThesaurus is a web-based system designed to map a comprehensive collection of protein and gene names to protein entries in the UniProt Knowledgebase. Currently covering more than two million proteins, BioThesaurus consists of over 2.8 million names extracted from multiple molecular biological databases according to the database cross-references in iProClass. The BioThesaurus web site allows the retrieval of synonymous names of given protein entries and the identification of protein entries sharing the same names. AVAILABILITY: BioThesaurus is accessible for online searching at http://pir.georgetown.edu/iprolink/biothesaurus

Animals↗

PromoterExplorer: an effective promoter identification method based on the AdaBoost algorithm.

MOTIVATION: Promoter prediction is important for the analysis of gene regulations. Although a number of promoter prediction algorithms have been reported in literature, significant improvement in prediction accuracy remains a challenge. In this paper, an effective promoter identification algorithm, which is called PromoterExplorer, is proposed. In our approach, we analyze the different roles of various features, that is, local distribution of pentamers, positional CpG island features and digitized DNA sequence, and then combine them to build a high-dimensional input vector. A cascade AdaBoost-based learning procedure is adopted to select the most 'informative' or 'discriminating' features to build a sequence of weak classifiers, which are combined to form a strong classifier so as to achieve a better performance. The cascade structure used for identification can also reduce the false positive. RESULTS: PromoterExplorer is tested based on large-scale DNA sequences from different databases, including the EPD, DBTSS, GenBank and human chromosome 22. Experimental results show that consistent and promising performance can be achieved.

Algorithms↗

Testing the hypothesis of recent population expansions in nematode parasites of human-associated hosts.

It has been predicted that parasites of human-associated organisms (eg humans, domestic pets, farm animals, agricultural and silvicultural plants) are more likely to show rapid recent population expansions than are parasites of other hosts. Here, we directly test the generality of this demographic prediction for species of parasitic nematodes that currently have mitochondrial sequence data available in the literature or the public-access genetic databases. Of the 23 host/parasite combinations analysed, there are seven human-associated parasite species with expanding populations and three without, and there are three non-human-associated parasite species with expanding populations and 10 without. This statistically significant pattern confirms the prediction. However, it is likely that the situation is more complicated than the simple hypothesis test suggests, and those species that do not fit the predicted general pattern provide interesting insights into other evolutionary processes that influence the historical population genetics of host-parasite relationships. These processes include the effects of postglacial migrations, evolutionary relationships and possibly life-history characteristics. Furthermore, the analysis highlights the limitations of this form of bioinformatic data-mining, in comparison to controlled experimental hypothesis tests.

Animals↗

Optimal word sizes for dissimilarity measures and estimation of the degree of dissimilarity between DNA sequences.

MOTIVATION: Several measures of DNA sequence dissimilarity have been developed. The purpose of this paper is 3-fold. Firstly, we compare the performance of several word-based or alignment-based methods. Secondly, we give a general guideline for choosing the window size and determining the optimal word sizes for several word-based measures at different window sizes. Thirdly, we use a large-scale simulation method to simulate data from the distribution of SK-LD (symmetric Kullback-Leibler discrepancy). These simulated data can be used to estimate the degree of dissimilarity beta between any pair of DNA sequences. RESULTS: Our study shows (1) for whole sequence similiarity/dissimilarity identification the window size taken should be as large as possible, but probably not >3000, as restricted by CPU time in practice, (2) for each measure the optimal word size increases with window size, (3) when the optimal word size is used, SK-LD performance is superior in both simulation and real data analysis, (4) the estimate beta of beta based on SK-LD can be used to filter out quickly a large number of dissimilar sequences and speed alignment-based database search for similar sequences and (5) beta is also applicable in local similarity comparison situations. For example, it can help in selecting oligo probes with high specificity and, therefore, has potential in probe design for microarrays. AVAILABILITY: The algorithm SK-LD, estimate beta and simulation software are implemented in MATLAB code, and are available at http://www.stat.ncku.edu.tw/tjwu

Algorithms↗

Identification of the proliferation/differentiation switch in the cellular network of multicellular organisms.

The protein-protein interaction networks, or interactome networks, have been shown to have dynamic modular structures, yet the functional connections between and among the modules are less well understood. Here, using a new pipeline to integrate the interactome and the transcriptome, we identified a pair of transcriptionally anticorrelated modules, each consisting of hundreds of genes in multicellular interactome networks across different individuals and populations. The two modules are associated with cellular proliferation and differentiation, respectively. The proliferation module is conserved among eukaryotic organisms, whereas the differentiation module is specific to multicellular organisms. Upon differentiation of various tissues and cell lines from different organisms, the expression of the proliferation module is more uniformly suppressed, while the differentiation module is upregulated in a tissue- and species-specific manner. Our results indicate that even at the tissue and organism levels, proliferation and differentiation modules may correspond to two alternative states of the molecular network and may reflect a universal symbiotic relationship in a multicellular organism. Our analyses further predict that the proteins mediating the interactions between these modules may serve as modulators at the proliferation/differentiation switch.

Animals↗

A picture of gene sampling/expression in model organisms using ESTs and KOG proteins.

The expressed sequence tag (EST) is an instrument of gene discovery. When available in large numbers, ESTs may be used to estimate gene expression. We analyzed gene expression by EST sampling, using the KOG database, which includes 24,154 proteins from Arabidopsis thaliana (Ath), 17,101 from Caenorhabditis elegans (Cel), 10,517 from Drosophila melanogaster (Dme), and 26,324 from Homo sapiens (Hsa), and 178,538 ESTs for Ath, 215,200 for Cel, 261,404 for Dme, and 1,941,556 for Hsa. BLAST similarity searches were performed to assign KOG annotation to all ESTs. We determined the amount of gene sampling or expression dedicated to each KOG functional category by each model organism. We found that the 25% most-expressed genes are frequently shared among these organisms. The KOG protein classification allowed the EST sampling calculation throughout the glycolysis pathway. We calculated the KOG cluster coverage and inferred that 50 to 80 K ESTs would efficiently cover 80-85% of the KOG database clusters in a transcriptome project. Since KOG is a database biased towards housekeeping genes, this is probably the number of ESTs needed to include the more commonly expressed genes in these organisms. We also examined a still unaddressed question: what is the minimum number of ESTs that should be produced in a transcriptome project?

Animals↗

Genetic variability at the D1S80 minisatellite: predominance of allele 18 among some indian populations.

BACKGROUND: Hypervariable minisatellites are considered as useful genetic markers in population studies because they are highly polymorphic, multiallelic and co-dominant in nature. The D1S80 minisatellite is one of the well studied markers, and has been used for differentiating population groups of various geographic, linguistic, cultural and genetic origins. OBJECTIVE: The present study reports the genetic variation observed at the D1S80 minisatellite among seven anthropologically distinct ethnic groups from Kerala state in south India and is compared with other reported Indian and world populations. SUBJECTS AND METHODS: DNA was isolated from the peripheral blood samples of 282 random, normal and healthy volunteers, PCR amplified and electrophoresed on 4% PAGE followed by silver staining. RESULTS: A total of 22 alleles (14-39 repeats) were detected with high heterozygosity (0.63-0.84) and Polymorphic Information Content (PIC) values (0.63-0.83). Allele 18 was the predominant allele, except in Ezhavas. The comparison of allele frequency data with world populations including other studied Indian ethnic groups has revealed that the majority of Indian populations possessed allele 18 as the predominant allele. In contrast, allele 24 was reported to be the predominant allele worldwide with a few exceptions. CONCLUSIONS: This study at the D1S80 minisatellite on seven ethnic groups will provide useful information for the Indian population genetic database. However, the most important observation was the predominance of allele 18 among the majority of Indian ethnic groups. The reason is not clear yet and thus further studies on Indian ethnic groups from different regions are necessary to find out the importance of allele 18 as the predominant allele in Indian population.

Adult↗

Influence of ADRB2 variants on bronchodilator response and asthma control in a mixed population.

OBJECTIVE: Given that b2 agonists constitute the primary treatment for asthma and that treatment response varies as a result of polymorphisms in the ADRB2 gene, we sought to investigate the associations between ADRB2 gene variants and bronchodilator response (BDR) in asthma patients. METHODS: A genetic database comprising 813 individuals was analyzed for variants in the ADRB2 gene. A longitudinal analysis of severe asthma patients was performed to evaluate changes in BDR over time. RESULTS: The rs1042713, rs1042714, and rs1042717 variants were associated with age-related changes in BDR in patients with severe asthma. The G allele (rs1042714) and the A allele (rs1042717) were associated with uncontrolled asthma, with carriers of the G46/G79/A252 alleles showing a higher risk of difficult-to-control asthma. Notably, no association was found between these variants and ADRB2 expression levels. CONCLUSIONS: Our findings suggest that a genetic panel including ADRB2 variants, as well as age-related differences in BDR, is a useful complementary tool in asthma management.

Humans↗