Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Databases, Genetic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Herpes glycoprotein gL is distantly related to chemokine receptor ligands.

Glycoprotein L (gL) is one of the critical proteins involved in transmission of Herpesviridae. We applied the methodology of protein structure prediction to shed a light on the so far unknown molecular mechanism of its action. Here we show that gL forms a chemokine-like protein. Alphaherpesvirinae gL as well as CMV functional homolog (UL130) create a novel CX chemokine-like protein, while Gammaherpesvirinae gL (HHV8 and EBV) adopt a regular CC beta-chemokine fold. We conclude that gL may interact with specific cellular chemokine receptors during the invasion of Herpesviridae. The proposed mechanism has a potential impact on future development of novel therapeutic and prophylactic strategies.

Amino Acid Motifs↗

A proteome-wide analysis of domain architectures of prokaryotic single-spanning transmembrane proteins.

We performed a proteome-wide survey of the domain architectures in single-spanning transmembrane (TM) proteins (single-spannings) from 87 sequenced prokaryotic (Bacterial and Archaean) genomes by assigning Pfam domains to their N-tail and C-tail loops. Out of 14,625 single-spannings, 3,516 sequences have at least one domain assigned, and no domains were assigned to 7,850, with the remaining 3,259 with less reliable assignment. In the domain-assigned sequences, 3116 sequences are with at most two domains, and the other 400 sequences with more than two. The assigned domains distribute over 651 Pfam families, which account for 11.4% of the total Pfam-A families. Among the 651 families are mostly soluble-protein-originated ones, but only 21 families are unique to TM proteins. The occurrence frequency of the individual domain families follows a power-law, that is, 264 families occur only once, 106 just twice, and the families appeared more than 30 times are counted by only 39. It is found that the great majority of the sequences having one or two domains are of the type II topology with the C-tail loop containing domains on it. On the contrary, the N-tail loop of the same type topology seldom carries domains. Importantly, the assigned domains are always found on the tail loops longer than 60 residues, even for the small domains with less than 30 residues. There are still as many as 5,800 sequences without assigned domains in spite of having at least one long tail, on which no less than 1,000 novel domain families are expected most likely to lie concealed unknown yet. We also investigated the domain arrangement preference and the domain family combination patterns in 'singlets' (single-spannings with one assigned domain) and 'doublets' (with two domains).

Computational Biology↗

Involvement of some large immunophilins and their ligands in the protection and regeneration of neurons: a hypothetical mode of action.

The powerful immunosuppressive drugs such as FK506 and its derivatives induce some regeneration and protection of neurons from ischaemic brain injury and some other neurological disorders. The drugs form complexes with diverse FKBPs but apparently the FKBP52/FK506 complex was shown to be involved in the protection and regeneration of neurons. We used several different sequence attributes in searching diverse genomic databases for similar motifs as those present in the FKBPs. A Fortran library of algorithms (Par_Seq) has been designed and used in searching for the similarity of sequence motifs extracted from the multiple sequence alignments of diverse groups of proteins (query motifs) and the target motifs which are encoded in various genomes. The following sequence attributes were used in the establishment of the degree of convergence between: (A) amino acid (AA) sequence similarity (ID) of the query/target motifs and (B) their: (1) AA composition (AAC); (2) hydrophobicity (HI); (3) Jensen-Shannon entropy; and (4) AA propensity to form a particular secondary structure. The sequence hallmark of two different groups of peptidylprolyl cis/trans isomerases (PPIases), namely tetratricopetide repeat (TPR) motifs, which are present in the heat-shock cyclophilins and in the large FK506-binding proteins (FKBPs) were used to search various genomic databases. The Par_Seq algorithm has revealed that the TPR motifs have similar sequence attributes as a number of hydrophobic sequence segments of functionally unrelated membrane proteins, including some of the TMs from diverse G protein-coupled receptors (GPCRs). It is proposed that binding of the FKBP52/FK506 complex to the membranes via the TPR motifs and its interaction with some membrane proteins could be in part responsible for some neuro-regeneration and neuro-protection of the brain during some ischaemia-induced stresses.

Algorithms↗

Bridging the gap between medical and bioinformatics: an ontological case study in colon carcinoma.

Ontological principles are needed in order to bridge the gap between medical and biological information in a robust and computable fashion. This is essential in order to draw inferences across the levels of granularity which span medicine and biology, an example of which include the understanding of the roles of tumor markers in the development and progress of carcinoma. Such information integration is also important for the integration of genomics information with the information contained in the electronic patient records in such a way that real time conclusions can be drawn. In this paper, we describe a large multi-granular datasource built by using ontological principles and focusing on the case of colon carcinoma.

Colonic Neoplasms↗

Aspergillus nidulans swoK encodes an RNA binding protein that is important for cell polarity.

The Aspergillus nidulans swoK1 mutant is defective in polarity maintenance when grown at restrictive temperature (38 degrees C). Upon germination, the mutant extends a primary germ tube that swells to an enlarged, non-uniform cell with pronounced wall thickenings. The mutant is fully restored to wild-type growth when transformed with a plasmid containing the AN5802.2 ORF as designated in The Broad Institute A. nidulans sequence database. Genetic mapping places swoK in the same region of chromosome I, as that occupied by An5802.2 on the physical map. swoK is predicted to encode a protein that contains an N-terminal RRM (RNA Recognition Motif) and a highly repetitive C-terminus with numerous RD/DR and RS/SR dipeptides. We hypothesize that SwoK participates in one of the known functions of SR proteins (those that contain SR/RS repeats): mRNA maturation through the spliceosome and or transport of mRNAs out of the nucleus to sites of protein translation.

Amino Acid Sequence↗

Insights into recombination from population genetic variation.

Patterns of genetic variation in natural populations are shaped by, and hence carry valuable information about, the underlying recombination process. In the past five years, the increasing availability of large-scale population genetic data on dense sets of markers, coupled with advances in statistical methods for extracting information from these data, have led to several important advances in our understanding of the recombination process in humans. These advances include the identification of large numbers of 'hotspots', where recombination appears to take place considerably more frequently than in the surrounding sequence, and the identification of DNA sequence motifs that are associated with the locations of these hotspots.

Animals↗

Evolution of metabolic networks by gain and loss of enzymatic reaction in eukaryotes.

The metabolic network is composed of enzymatic reactions (ERs) in which one or more enzymes catalyze the reaction of pertinent substrates. Since metabolism is a basal system for maintaining life of all organisms, any change in the metabolic networks must greatly affect organismic evolution. The aim of this study is to examine how often gains and losses of ER have occurred during the evolution of metabolic networks in eukaryotes and how these evolutionary events have affected phenotypic traits of organisms. In this study, we conducted comparative studies of 751 ERs in the metabolic networks of 6 eukaryotic species whose complete genome sequences were determined. As a result, we found that a total of 804 gains and losses of ERs had occurred in the evolutionary diversification of metabolic networks in different lineages. Moreover, the vertebrate lineage, after the separation from Drosophila melanogaster, showed a remarkable increase in the number of ER gains compared with ER losses. In particular, 41% of the ER gains were predominantly involved with lipids and complex lipid metabolism. Because some products of these two metabolisms function as hormones, we concluded that the ER gain of these two metabolisms accelerated the development of hormonal signal transduction for the elaborate regulation of physiological systems during vertebrate evolution.

Amino Acid Sequence↗

QSPR models for the prediction of apparent volume of distribution.

An estimate of volume of distribution (V(d)) is of paramount importance both in drug choice as well as maintenance and loading dose calculations in therapeutics. It can also be used in the prediction of drug biological half life. This study employs quantitative structure-pharmacokinetic relationship (QSPR) techniques for the prediction of volume of distribution. Values of V(d) for 129 drugs were collated from the literature. Structural descriptors consisted of partitioning, quantum mechanical, molecular mechanical, and connectivity parameters calculated by specialized software and pK(a) values obtained from ACD labs/log D database. Genetic algorithm and stepwise regression analyses were used for variable selection and model development. Models were validated using a leave-many-out procedure. QSPR analyses resulted in a number of significant models for acidic and basic drugs separately, and for all the drugs. Validation studies showed that mean fold error of predictions for the selected models were between 1.79 and 2.17. Although separate QSPR models for acids and bases resulted in lower prediction errors than models for all the drugs, the external validation study showed a limited applicability for the equation obtained for acids. Therefore, the universal model that requires only calculated structural descriptors was recommended. The QSPR model is able to predict the volume of distribution of drugs belonging to different chemical classes with a prediction error similar to that of the other more complicated prediction methods including the commonly practiced interspecies scaling. The structural descriptors in the model can be interpreted based on the known mechanisms of distribution and the molecular structures of the drugs.

Algorithms↗

Use of morphological analysis in protein name recognition.

Protein name recognition aims to detect each and every protein names appearing in a PubMed abstract. The task is not simple, as the graphic word boundary (space separator) assumed in conventional preprocessing does not necessarily coincide with the protein name boundary. Such boundary disagreement caused by tokenization ambiguity has usually been ignored in conventional preprocessing of general English. In this paper, we argue that boundary disagreement poses serious limitations in biomedical English text processing, not to mention protein name recognition. Our key idea for dealing with the boundary disagreement is to apply techniques used in Japanese morphological analysis where there are no word boundaries. Having evaluated the proposed method with GENIA corpus 3.02, we obtain F-measure of 69.01 on a strict criterion and 79.32 on a relaxed criterion. The result is comparable to other published work in protein name recognition, without resorting to manually prepared ad hoc feature engineering. Further, compared to the conventional preprocessing, the use of morphological analysis as preprocessing improves the performance of protein name recognition and reduces the execution time.

Abstracting and Indexing↗

Term identification in the biomedical literature.

Sophisticated information technologies are needed for effective data acquisition and integration from a growing body of the biomedical literature. Successful term identification is key to getting access to the stored literature information, as it is the terms (and their relationships) that convey knowledge across scientific articles. Due to the complexities of a dynamically changing biomedical terminology, term identification has been recognized as the current bottleneck in text mining, and--as a consequence--has become an important research topic both in natural language processing and biomedical communities. This article overviews state-of-the-art approaches in term identification. The process of identifying terms is analysed through three steps: term recognition, term classification, and term mapping. For each step, main approaches and general trends, along with the major problems, are discussed. By assessing previous work in context of the overall term identification process, the review also tries to delineate needs for future work in the field.

Abbreviations as Topic↗

Enhancing performance of protein and gene name recognizers with filtering and integration strategies.

Named entity (NE) recognition is a fundamental task in biological relationship mining. This paper considers protein/gene collocates extracted from biological corpora as restrictions to enhance the precision rate of protein/gene name recognition. In addition, we integrate the results of multiple NE recognizers to improve the recall rates. Yapex and KeX, and ABGene and Idgene are taken as examples of protein and gene name recognizers, respectively. The precision of Yapex increases from 70.90 to 85.84% at the low expense of the recall rate (i.e., it only decreases 2.44%) when collocates are incorporated. When both filtering and integration strategies are employed together, the Yapex-based integration with KeX shows good performance, i.e., the F-score increases by 7.83% compared to the pure Yapex method. The results of gene recognition show the same tendency. The ABGene-based integration with Idgene shows a 10.18% F-score increase compared to the pure ABGene method. These successful methodologies can be easily extended to other name finders in biological documents.

Abstracting and Indexing↗

Evaluation of techniques for increasing recall in a dictionary approach to gene and protein name identification.

Gene and protein name identification in text requires a dictionary approach to relate synonyms to the same gene or protein, and to link names to external databases. However, existing dictionaries are incomplete. We investigate two complementary methods for automatic generation of a comprehensive dictionary: combination of information from existing gene and protein databases and rule-based generation of spelling variations. Both methods have been reported in literature before, but have hitherto not been combined and evaluated systematically. We combined gene and protein names from several existing databases of four different organisms. The combined dictionaries showed a substantial increase in recall on three different test sets, as compared to any single database. Application of 23 spelling variation rules to the combined dictionaries further increased recall. However, many rules appeared to have no effect and some appear to have a detrimental effect on precision.

Abstracting and Indexing↗

Global gene expression profiling and cluster analysis in Xenopus laevis.

We have undertaken a large-scale microarray gene expression analysis using cDNAs corresponding to 21,000 Xenopus laevis ESTs. mRNAs from 37 samples, including embryos and adult organs, were profiled. Cluster analysis of embryos of different stages was carried out and revealed expected affinities between gastrulae and neurulae, as well as between advanced neurulae and tadpoles, while egg and feeding larvae were clearly separated. Cluster analysis of adult organs showed some unexpected tissue-relatedness, e.g. kidney is more related to endodermal than to mesodermal tissues and the brain is separated from other neuroectodermal derivatives. Cluster analysis of genes revealed major phases of co-ordinate gene expression between egg and adult stages. During the maternal-early embryonic phase, genes maintaining a rapidly dividing cell state are predominantly expressed (cell cycle regulators, chromatin proteins). Genes involved in protein biosynthesis are progressively induced from mid-embryogenesis onwards. The larval-adult phase is characterised by expression of genes involved in metabolism and terminal differentiation. Thirteen potential synexpression groups were identified, which encompass components of diverse molecular processes or supra-molecular structures, including chromatin, RNA processing and nucleolar function, cell cycle, respiratory chain/Krebs cycle, protein biosynthesis, endoplasmic reticulum, vesicle transport, synaptic vesicle, microtubule, intermediate filament, epithelial proteins and collagen. Data filtering identified genes with potential stage-, region- and organ-specific expression. The dataset was assembled in the iChip microarray database, , which allows user-defined queries. The study provides insights into the higher order of vertebrate gene expression, identifies synexpression groups and marker genes, and makes predictions for the biological role of numerous uncharacterized genes.

Animals↗

Evolution of vertebrate haemoglobins: Histidine side chains, specific buffer value and Bohr effect.

This review highlights the use of analytical tools, recently developed in the comparative method of evolutionary biology, for the study of haemoglobin (Hb) adaptation. It focuses on the functional consequences of a previously largely ignored structural feature of Hb, namely the degree and positional specificity of histidine (His) substitution in Hb chains. The importance of His side chains for hydrogen ion buffering, blood CO(2) transport capacity and the molecular mechanism of the Bohr effect in vertebrate Hbs is discussed. Using phylogenetically independent contrasts, a significant correlation between the specific buffer value of Hb and the number of predicted physiological buffer groups from Hb sequence data is shown. In a new result, the evolution of the number of physiological buffer groups in 77 vertebrate species is reconstructed on a phylogenetic tree. The analysis predicts that teleost fishes, passeriform birds and some snakes have independently evolved a much-reduced specific buffer value of Hb, possibly for enhancing the efficiency of an acid load to change oxygen affinity via the Bohr effect. This analysis demonstrates how in comparative physiology analysis of genetic databases in an evolutionary framework can identify candidate species for further experimental in vitro and whole animal studies.

Animals↗

Integrating 'omic' information: a bridge between genomics and systems biology.

The availability of genome sequences for several organisms, including humans, and the resulting first-approximation lists of genes, have allowed a transition from molecular biology to 'modular biology'. In modular biology, biological processes of interest, or modules, are studied as complex systems of functionally interacting macromolecules. Functional genomic and proteomic ('omic') approaches can be helpful to accelerate the identification of the genes and gene products involved in particular modules, and to describe the functional relationships between them. However, the data emerging from individual omic approaches should be viewed with caution because of the occurrence of false-negative and false-positive results and because single annotations are not sufficient for an understanding of gene function. To increase the reliability of gene function annotation, multiple independent datasets need to be integrated. Here, we review the recent development of strategies for such integration and we argue that these will be important for a systems approach to modular biology.

Computational Biology↗

Long-read Sequences Mapped to a Complete Reference Genome Uncover Uncaptured Structural Variants across the Beta-globin Cluster in Africans with Sickle Cell Disease.

African genomes are marked by extensive complexity in the number and distribution of variants, yet remain under-represented in genetic databases and the human reference genome. This gap in representation limits the broad application of genomic medicine. Sickle cell disease (SCD) - one of the most common monogenic diseases - has its highest prevalence in Africa, and variation in disease severity has consistently been linked to the beta-globin locus, including levels of fetal hemoglobin (HbF). Modulation of HbF is central to current SCD gene therapies; however, the inherent complexity and variation at the locus in African genomes presents a challenge to translating these advances to Africa. Here, we align long-read single molecule sequences (LRS) targeted to the beta-globin region to the hg38 and T2T-CHM13v2 genome references in 40 individuals with SCD, predominantly recruited from three African countries. We demonstrate that the expanded T2T-CHM13v2 reference sequence at this locus reduces Structural Variant (SV) calls by 70% and uncovers uncaptured single nucleotide variants (SNVs). Across the cluster we report 343 SVs and 196 SNVs that have not been previously reported, including in LRS data from the All of Us project. By including African populations from ethnolinguistic groups that have not been previously surveyed we improve variant resolution and bolster evidence for observed variation. Finally, we identify a common ∼4kb insertion locus overlapping the HBB promoter among individuals with high HbF. These results demonstrate the utility of combining a comprehensive reference genome with LRS in African populations to uncover genomic variation at disease-associated loci.

SNV↗

Informatics for protein identification by mass spectrometry.

High throughput protein analysis (i.e., proteomics) first became possible when sensitive peptide mass mapping techniques were developed, thereby allowing for the possibility of identifying and cataloging most 2D gel electrophoresis spots. Shortly thereafter a few groups pioneered the idea of identifying proteins by using peptide tandem mass spectra to search protein sequence databases. Hence, it became possible to identify proteins from very complex mixtures. One drawback to these latter techniques is that it is not entirely straightforward to make matches using tandem mass spectra of peptides that are modified or have sequences that differ slightly from what is present in the sequence database that is being searched. This has been part of the motivation behind automated de novo sequencing programs that attempt to derive a peptide sequence regardless of its presence in a sequence database. The sequence candidates thus generated are then subjected to homology-based database search programs (e.g., BLAST or FASTA). These homology search programs, however, were not developed with mass spectrometry in mind, and it became necessary to make minor modifications such that mass spectrometric ambiguities can be taken into account when comparing query and database sequences. Finally, this review will discuss the important issue of validating protein identifications. All of the search programs will produce a top ranked answer; however, only the credulous are willing to accept them carte blanche.

Amino Acid Sequence↗

Ultrasound findings and multiple marker screening in trisomy 18.

OBJECTIVE: To compare detection of trisomy 18 in the second trimester by ultrasound and multiple-marker testing. METHODS: A computerized genetics database was used to identify fetuses of 14-22 weeks' gestation who had comprehensive ultrasound examinations, multiple-marker screening tests (alpha-fetoprotein [AFP]), hCG, unconjugated estriol [E3], and trisomy 18 karyotype. A positive trisomy 18 screen was defined as AFP up to 0.75 multiples of the median (MoM), hCG up to 0.55 MoM, and unconjugated E3 up to 0.60 MoM. A risk of at least 1:190 defined a positive Down syndrome screen. Ultrasound abnormalities were diagnosed prospectively and were confirmed later by retrospective review of sonographic images. RESULTS: From 1988-1997, 30 trisomy 18 fetuses who had comprehensive ultrasounds and multiple-marker testing were identified. Twenty-one (70%) had abnormalities detected by ultrasound, of which the most common isolated finding was choroid plexus cyst. Eleven fetuses (37%) had positive trisomy 18 screens, and two had positive Down syndrome screens, for a total of 13 of 30 (43%) fetuses with positive multiple-marker screening tests. CONCLUSION: We found that ultrasound was more likely to be abnormal than multiple-marker screening tests in fetuses with trisomy 18 (70%) (95% confidence interval [CI] 54, 86 versus 43% CI 25, 61). However, combining the two testing methods yielded the highest detection rate (80% [CI 66%, 94%]).

Adult↗