Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Spectral karyotyping and fluorescence in situ hybridization detect novel chromosomal aberrations, a recurring involvement of chromosome 21 and amplification of the MYC oncogene in acute myeloid leukaemia M2.

Recurring chromosomal aberrations are of aetiological, diagnostic, prognostic and therapeutic importance in acute myeloid leukaemia (AML). However, aberrations are detected in only two thirds of AML cases at diagnosis and recurrent balanced translocations in only 50%. Spectral karyotyping (SKY) enables simultaneous visualization of all human chromosomes in different colours, facilitating the comprehensive evaluation of chromosomal abnormalities. Therefore, SKY was used to characterize 37 cases of newly diagnosed AML-M2, previously analysed using G-banding. In 15/23 patients it was possible to obtain metaphases from viably frozen cells; in 22 additional cases, fixed-cell suspensions were used. Of the 70 chromosomal aberrations identified by SKY, 30 aberrations were detected for the first time, 18 aberrations were redefined and 22 were confirmed. SKY detected two reciprocal translocations, t(X;3) and t(11;19). In five cases, eight structural aberrations resulted in partial gains of chromosome 21, six of which were undetected by G-banding. In 4/5 cases, these resulted in copy number increases for AML1. Amplification of MYC was detected in three cases. Using SKY and FISH, clonal aberrations were identified in 5/18 cases with a presumed normal karyotype; 3/5 aberrations were of known unfavourable prognostic significance. Karyotypes were entered into a custom-designed SKY database, which will be integrated with other cytogenetic and genomic databases.

Chromosome Aberrations↗

In silico cloning of mouse Muc5b gene and upregulation of its expression in mouse asthma model.

Using a BLAST-searching approach, we identified a mouse expressed sequence tag (EST) clone (AA038672) showing great similarity to the 3' end of the human MUC5B gene. The clone was named "3pmmuc5b-1" after complete nucleotide sequencing (Genbank Accession, AF369933). A subsequent search of the mouse genome database with the 3pmmuc5b-1 sequence identified two overlapping genomic clones (AC020817 and AC020794) that contained the sequence of both 3pmmuc5b-1 and the mouse Muc5ac gene. Like their human homologs, the genomic order of the mouse Muc genes is 5'-Muc5ac-Muc5b-3'. These results suggest that the newly identified EST clone, 3pmmuc5b-1, is part of the 3' portion of the mouse Muc5b gene. In situ hybridization demonstrated that this putative mouse Muc5b message was expressed in a restricted manner in the sublingual gland region of the tongue and the submucosal gland region of the mouse trachea in a normal animal. However, the gene expression was greatly enhanced in airway surface epithelium and the submucosal gland region in ovalbumin-induced asthmatic mice. These results were consistent with previous studies of human airway tissues. We therefore conclude that this newly cloned mouse Muc5b gene could be used as a marker for studying aberrant mucin gene expression in mouse models of various airway diseases.

Animals↗

Purification, characterization and sequence analysis of Omp50,a new porin isolated from Campylobacter jejuni.

A novel pore-forming protein identified in Campylobacter was purified by ion-exchange chromatography and named Omp50 according to both its molecular mass and its outer membrane localization. We observed a pore-forming ability of Omp50 after re-incorporation into artificial membranes. The protein induced cation-selective channels with major conductance values of 50-60 pS in 1 M NaCl. N-terminal sequencing allowed us to identify the predicted coding sequence Cj1170c from the Campylobacter jejuni genome database as the corresponding gene in the NCTC 11168 genome sequence. The gene, designated omp50, consists of a 1425 bp open reading frame encoding a deduced 453-amino acid protein with a calculated pI of 5.81 and a molecular mass of 51169.2 Da. The protein possessed a 20-amino acid leader sequence. No significant similarity was found between Omp50 and porin protein sequences already determined. Moreover, the protein showed only weak sequence identity with the major outer-membrane protein (MOMP) of Campylobacter, correlating with the absence of antigenic cross-reactivity between these two proteins. Omp50 is expressed in C. jejuni and Campylobacter lari but not in Campylobacter coli. The gene, however, was detected in all three species by PCR. According to its conformation and functional properties, the protein would belong to the family of outer-membrane monomeric porins.

Amino Acid Sequence↗

Bioinformatics in medical practice: what is necessary for a hospital?

Building bioinformatic facilities for a university hospital is pretty similar to using standardized building blocks to construct a house. Starting with the intention to built a dwelling house, a factory or just a shelter the architect draws a construction plan and determines the material to be used. In general, the building is then constructed by the workmen following exactly the plan. However, for particular reasons, minor alterations may be needed to improve the construction of the building. Here we use the metaphor of constructing a "bio-informatics building" to describe the steps needed to support the daily tasks of a university hospital medical microbiology department which uses genomic methods quite extensively for pathogen identification. Today the Giessen "bioinformatics building" is not yet complete but we have been able to lay solid foundations and erect the ground floor which is functional already. Using a combination of standard tools, internet accessible genomic databases and some own software tools we can support genome sequencing from the raw sequence to pathogen identification.

Computational Biology↗

SMS: sequence, motif and structure--a database on the structural rigidity of peptide fragments in non-redundant proteins.

Structure prediction methods aim to identify the relationship between the amino acid sequence of an unknown protein and information comprised in databases of known protein structures. Towards this end, we created a database by combining the amino acid sequences and the corresponding three-dimensional atomic coordinates for all the 25% non-redundant protein chains available in the Protein Data Bank. It contains information about the peptide fragments that are 5 to 10 residues long. In addition, options are provided for the users to visualize the individual motifs and the superposed fragments in the client machine. Further, useful functionalities areprovided to look for similar sequence motifs in all the sequence databases like PDB, 90% non-redundant protein chains, Genome database, PIR and Swiss-Prot. The database is being updated at regular intervals and the same can be accessed over the World Wide Web interface at the following URL: http://pranag.physics.iisc.ernet.in/sms/.

Amino Acid Sequence↗

Searching genomes for noncoding RNA using FastR.

The discovery of novel noncoding RNAs has been among the most exciting recent developments in biology. It has been hypothesized that there is, in fact, an abundance of functional noncoding RNAs (ncRNAs) with various catalytic and regulatory functions. However, the inherent signal for ncRNA is weaker than the signal for protein coding genes, making these harder to identify. We consider the following problem: Given an RNA sequence with a known secondary structure, efficiently detect all structural homologs in a genomic database by computing the sequence and structure similarity to the query. Our approach, based on structural filters that eliminate a large portion of the database while retaining the true homologs, allows us to search a typical bacterial genome in minutes on a standard PC. The results are two orders of magnitude better than the currently available software for the problem. We applied FastR to the discovery of novel riboswitches, which are a class of RNA domains found in the untranslated regions. They are of interest because they regulate metabolite synthesis by directly binding metabolites. We searched all available eubacterial and archaeal genomes for riboswitches from purine, lysine, thiamin, and riboflavin subfamilies. Our results point to a number of novel candidates for each of these subfamilies and include genomes that were not known to contain riboswitches.

Algorithms↗

Bioinformatic analysis of the human mu opioid receptor (OPRM1) splice and polymorphic variants.

Mu opioid receptor (OPRM1), a member of the G-protein coupled receptor superfamily, mediates the analgesic and euphoric effects of opioid drugs. The sequences of OPRM1 cDNA and reported splice variants were used to search the public and Celera genomic databases. The matched sequences were analyzed to assemble an OPRM1 genomic contig. Human OPRM1 gene was estimated to span at least 90 kb in the chromosome 6q24-25 region. Four coding exons are separated by 3 introns. While intron 2 has only 773 bp, these databases for the first time provide the precise length of and other information about long introns 1 and 3, containing 50 and 27 kb, respectively. When a consensus exon/intron splice junction at the end of the coding exon 3 was not utilized, it may have resulted in continuous translation of the exon to yield the splice variant OPRM1A. The study did not identify human orthologs of other OPRM1 variants that had been reported for mouse OPRM1, although several proposed exons were found to be included in mouse genomic clones. Single nucleotide polymorphisms in the OPRM1 gene were also analyzed and summarized, which could provide potential polymorphic markers for molecular genetic studies.

Alternative Splicing↗

Gene maps and location databases.

A location database is defined in linear space by a vector of genetic and physical locations for each locus, which may be ordered by virtual sorting on composite location. This contrasts with an interval database defined in metric space, for which location must be inferred by list-processing from numbered intervals which are assigned different ordinals in different tables and overlap other intervals in many ways. A location database has been used for all well-studied experimental organisms. Principles for a human genome database may be derived from this experience.

Animals↗

National genomic projects in Asia and Africa: a review.

National genome projects (NGPs) are increasingly shaping precision medicine by improving representation of population-specific genetic diversity. This review compiles findings from NGPs across Asia and Africa, regions that remain underrepresented in global genomic databases despite their extensive demographic and genetic diversity. A total of 53 studies from 24 countries were identified to understand (1) the genomic approach utilized, (2) novel findings that have emerged, and (3) strategies for improving research in these regions. The NGPs implement population-based variome databases (20 NGPs), linear reference genome assemblies (8 NGPs), and graph-based pangenome assemblies (1 NGP). Novel variants ranged between 0.28% (China) and 19.6% (Iran), whereas rare variants accounted for up to 88.9% of the detected variants in the Chinese population. Each NGP documents its country's evolutionary and migration history, which impacts disease frequency and pharmacogenomic variants. Clinically, NGPs revealed strong population stratification in disease-associated and pharmacogenomic variants. For example, the GJB2 rs72474224 hearing-loss variant ranged from 13% in Vietnam and 12% in Hong Kong to 0.0894% in Turkey, while the VKORC1 rs9923231 pharmacogenomic variant reached 89.2% in Taiwan but was 20%-25% in European-related Russian subpopulations. These findings demonstrate that clinically relevant allele frequencies, pathogenicity assessments, and drug-response markers differ substantially across ancestries. This review highlights ongoing efforts and strategies to enhance the representativeness of genomic data through NGPs in Asia and Africa. We also suggest future directions for national projects, including integrating family-based studies, multi-omic data, and standardized pipelines to accelerate discovery and support the equitable implementation of precision medicine.

Humans↗

PlasmoDB: exploring genomics and post-genomics data of the malaria parasite, Plasmodium falciparum.

The recent completion of the genome sequence of Plasmodium falciparum 3D7 provides the foundation for genome-wide analysis of the parasite. In addition to DNA and gene sequence data, postgenomic methods including microarray-based transcript profiling and high-throughput proteomics are now accessible to Plasmodium researchers. The Plasmodium Genome database ( ) was developed to provide rapid and convenient access to the terabytes of genomic-scale data now being generated around the world. All data are available in a relational framework, permitting convenient downloading, browsing, and analysis. Combinatorial use of data analysis tools enables powerful data mining queries, such as combining gene and protein expression data to monitor changes through various life-cycle stages. Functional predictions can be used to explore potential targets for antimalarial drug development. This report outlines the use of PlasmoDB to examine redox-active functions in Plasmodium.

Animals↗

[Computer searching of new targets for antimicrobial drugs based on comparative analysis of genomes].

The progress in genome research allows to use genomic databases for drug discovery. The major interest consists in their usage for searching of new molecular targets for new drugs. It is especially important in the area of new antimicrobial drugs creation. In recent years, the applicability of genome analysis for solution of this problem was shown by different authors. We propose an approach for searching new targets for antimicrobial drugs based upon comparative analysis of genomes and samples from molecular databases. For each protein encoded by the target microorganism genome the conformance to a number of medico-biological and technological requirements was tested. The obtained evaluations were used for selection of potential targets. The approach was implemented in the original software GenMesh. It was successfully tested with known targets of drugs against tuberculosis when correct valuations were obtained. The attempt of new targets selection for design of new drugs against tuberculosis was done with reliable results.

Anti-Infective Agents↗

The human olfactory receptor gene family.

Humans perceive an immense variety of chemicals as having distinct odors. Odor perception initiates in the nose, where odorants are detected by a large family of olfactory receptors (ORs). ORs have diverse protein sequences but can be assigned to subfamilies on the basis of sequence relationships. Members of the same subfamily have related sequences and are likely to recognize structurally related odorants. To gain insight into the mechanisms underlying odor perception, we analyzed the human OR gene family. By searching the human genome database, we identified 339 intact OR genes and 297 OR pseudogenes. Determination of their genomic locations showed that OR genes are unevenly distributed among 51 different loci on 21 human chromosomes. Sequence comparisons showed that the human OR family is composed of 172 subfamilies. Types of odorant structures that may be recognized by some subfamilies were predicted by identifying subfamilies that contain ORs with known odor ligands or human homologs of such ORs. Analysis of the chromosomal locations of members of each OR subfamily revealed that most subfamilies are encoded by a single chromosomal locus. Moreover, many loci encode only one or a few subfamilies, suggesting that different parts of the genome may, to some extent, be involved in the detection of different types of odorant structural motifs.

Animals↗

Genome Annotation Transfer Utility (GATU): rapid annotation of viral genomes using a closely related reference genome.

BACKGROUND: Since DNA sequencing has become easier and cheaper, an increasing number of closely related viral genomes have been sequenced. However, many of these have been deposited in GenBank without annotations, severely limiting their value to researchers. While maintaining comprehensive genomic databases for a set of virus families at the Viral Bioinformatics Resource Center http://www.biovirus.org and Viral Bioinformatics - Canada http://www.virology.ca, we found that researchers were unnecessarily spending time annotating viral genomes that were close relatives of already annotated viruses. We have therefore designed and implemented a novel tool, Genome Annotation Transfer Utility (GATU), to transfer annotations from a previously annotated reference genome to a new target genome, thereby greatly reducing this laborious task. RESULTS: GATU transfers annotations from a reference genome to a closely related target genome, while still giving the user final control over which annotations should be included. GATU also detects open reading frames present in the target but not the reference genome and provides the user with a variety of bioinformatics tools to quickly determine if these ORFs should also be included in the annotation. After this process is complete, GATU saves the newly annotated genome as a GenBank, EMBL or XML-format file. The software is coded in Java and runs on a variety of computer platforms. Its user-friendly Graphical User Interface is specifically designed for users trained in the biological sciences. CONCLUSION: GATU greatly simplifies the initial stages of genome annotation by using a closely related genome as a reference. It is not intended to be a gene prediction tool or a "complete" annotation system, but we have found that it significantly reduces the time required for annotation of genes and mature peptides as well as helping to standardize gene names between related organisms by transferring reference genome annotations to the target genome. The program is freely available under the General Public License and can be accessed along with documentation and tutorial from http://www.virology.ca/gatu.

Amino Acid Sequence↗

Evolution of the mdg1 lineage of the Ty3/gypsy group of LTR retrotransposons in Anopheles gambiae.

So far, only a few retrovirus-like transposable elements (TEs) have been reported in Anopheles mosquitoes, although a large fraction of their genomes is made up of these middle repetitive sequences. By screening the A. gambiae genome databases, we have found 10 element families belonging to the mdg1 lineage of the Ty3/gypsy group of long terminal repeat (LTR) retrotransposons. These Anopheles families constitute a sister clade of the Drosophila representatives of this same lineage. According to the phylogenetic reconstruction of their open reading frame (ORF)2 enzymatic domains, the analysis of patterns of nucleotide substitution therein, and the estimation of the age of particular insertions, all these elements must have been active until quite recently, and some of them must be very young. On the other hand, the fact that all these element families are primarily composed of fragmentary copies (mostly solos) or full-length copies with inactivating mutations indicates that their turnover rate has been probably very low. Finally, incongruent phylogenies obtained from different regions of the elements strongly suggest that recombination has played a significant role in their evolutionary history.

Animals↗

The human and mouse repertoire of the adhesion family of G-protein-coupled receptors.

The adhesion G-protein-coupled receptors (GPCRs) (also termed LN-7TM or EGF-7TM receptors) are membrane-bound proteins with long N-termini containing multiple domains. Here, 2 new human adhesion-GPCRs, termed GPR133 and GPR144, have been found by searches done in the human genome databases. Both GPR133 and GPR144 have a GPS domain in their N-termini, while GPR144 also has a pentraxin domain. The phylogenetic analyses of the 2 new human receptors show that they group together without close relationship to the other adhesion-GPCRs. In addition to the human genes, mouse orthologues to those 2 and 15 other mouse orthologues to human were identified (GPR110, GPR111, GPR112, GPR113, GPR114, GPR115, GPR116, GPR123, GPR124, GPR125, GPR126, GPR128, LEC1, LEC2, and LEC3). Currently the total number of human adhesion-GPCRs is 33. The mouse and human sequences show a clear one-to-one relationship, with the exception of EMR2 and EMR3, which do not seem to have orthologues in mouse. EST expression charts for the entire repertoire of adhesion-GPCRs in human and mouse were established. Over 1600 ESTs were found for these receptors, showing widespread distribution in both central and peripheral tissues. The expression patterns are highly variable between different receptors, indicating that they participate in a number of physiological processes.

Animals↗

"Iceland Inc."?: On the ethics of commercial population genomics.

A detailed analysis of the Icelandic commercial population-wide genomics database project of deCODE Genetics was performed for the purpose of providing ethics insights into public/private efforts to develop genetic databases. This analysis examines the moral differences between the general case of governmental collection of medical data for public health purposes and the centralized collection planned in Iceland. Both the process of developing the database and its design vary in significant ways from typical government data collection and analysis activities. Because of these differences, the database may serve the interests of deCODE more than it serves the interests of the public, undermining the claim that presumed consent for this data collection and its proprietary use is ethical. We believe that there is an evolving consensus that informed consent of participants must be secured for population-based genetics databases and research. The Iceland model provides an informative counterexample that holds key ethics lessons for similar ventures.

Databases, Genetic↗

Privacy-preserving framework for genomic computations via multi-key homomorphic encryption.

MOTIVATION: The affordability of genome sequencing and the widespread availability of genomic data have opened up new medical possibilities. Nevertheless, they also raise significant concerns regarding privacy due to the sensitive information they encompass. These privacy implications act as barriers to medical research and data availability. Researchers have proposed privacy-preserving techniques to address this, with cryptography-based methods showing the most promise. However, existing cryptography-based designs lack (i) interoperability, (ii) scalability, (iii) a high degree of privacy (i.e. compromise one to have the other), or (iv) multiparty analyses support (as most existing schemes process genomic information of each party individually). Overcoming these limitations is essential to unlocking the full potential of genomic data while ensuring privacy and data utility. Further research and development are needed to advance privacy-preserving techniques in genomics, focusing on achieving interoperability and scalability, preserving data utility, and enabling secure multiparty computation. RESULTS: This study aims to overcome the limitations of current cryptography-based techniques by employing a multi-key homomorphic encryption scheme. By utilizing this scheme, we have developed a comprehensive protocol capable of conducting diverse genomic analyses. Our protocol facilitates interoperability among individual genome processing and enables multiparty tests, analyses of genomic databases, and operations involving multiple databases. Consequently, our approach represents an innovative advancement in secure genomic data processing, offering enhanced protection and privacy measures. AVAILABILITY AND IMPLEMENTATION: All associated code and documentation are available at https://github.com/farahpoor/smkhe.

Computer Security↗

Genome analysis and phylogenetic relationships between east, central and west African isolates of Yellow fever virus.

Yellow fever virus (YFV), a reemerging disease agent in Africa and South America, is the prototype member of the genus Flavivirus. Based on examination of the prM/M, E and 3' non-coding regions of the YFV genome, previous studies have identified seven genotypes of YFV, including the Angolan, east/central African and east African genotypes, which are highly divergent from the prototype strain Asibi. In this study, full genome analysis was used to expand upon these genetic relationships as well as on the very limited full genome database for YFV. This study was the first to investigate genomic sequences of YFV strains from east and central Africa (Angola71, Uganda48a and Ethiopia61b). All three viruses had genomes of 10 823 nt in length. Compared with the prototype strain Asibi (from west Africa) they were approximately 25 % divergent in nucleotide sequence and 7 % divergent in amino acid sequence. Comparison of multiple flaviviruses in the N-terminal region of NS4B showed that amino acid sequences were variable and that west African strains of YFV had an amino acid deletion at residue 21. Additionally, N-linked glycosylation sites were conserved between viral genotypes, while codon usage varied between strains.

Africa, Central↗