Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Databases, Genetic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

TaxMan: a taxonomic database manager.

BACKGROUND: Phylogenetic analysis of large, multiple-gene datasets, assembled from public sequence databases, is rapidly becoming a popular way to approach difficult phylogenetic problems. Supermatrices (concatenated multiple sequence alignments of multiple genes) can yield more phylogenetic signal than individual genes. However, manually assembling such datasets for a large taxonomic group is time-consuming and error-prone. Additionally, sequence curation, alignment and assessment of the results of phylogenetic analysis are made particularly difficult by the potential for a given gene in a given species to be unrepresented, or to be represented by multiple or partial sequences. We have developed a software package, TaxMan, that largely automates the processes of sequence acquisition, consensus building, alignment and taxon selection to facilitate this type of phylogenetic study. RESULTS: TaxMan uses freely available tools to allow rapid assembly, storage and analysis of large, aligned DNA and protein sequence datasets for user-defined sets of species and genes. The user provides GenBank format files and a list of gene names and synonyms for the loci to analyse. Sequences are extracted from the GenBank files on the basis of annotation and sequence similarity. Consensus sequences are built automatically. Alignment is carried out (where possible, at the protein level) and aligned sequences are stored in a database. TaxMan can automatically determine the best subset of taxa to examine phylogeny at a given taxonomic level. By using the stored aligned sequences, large concatenated multiple sequence alignments can be generated rapidly for a subset and output in analysis-ready file formats. Trees resulting from phylogenetic analysis can be stored and compared with a reference taxonomy. CONCLUSION: TaxMan allows rapid automated assembly of a multigene datasets of aligned sequences for large taxonomic groups. By extracting sequences on the basis of both annotation and BLAST similarity, it ensures that all available sequence data can be brought to bear on a phylogenetic problem, but remains fast enough to cope with many thousands of records. By automatically assisting in the selection of the best subset of taxa to address a particular phylogenetic problem, TaxMan greatly speeds up the process of generating multiple sequence alignments for phylogenetic analysis. Our results indicate that an automated phylogenetic workbench can be a useful tool when correctly guided by user knowledge.

Database Management Systems↗

BIOZON: a system for unification, management and analysis of heterogeneous biological data.

BACKGROUND: Integration of heterogeneous data types is a challenging problem, especially in biology, where the number of databases and data types increase rapidly. Amongst the problems that one has to face are integrity, consistency, redundancy, connectivity, expressiveness and updatability. DESCRIPTION: Here we present a system (Biozon) that addresses these problems, and offers biologists a new knowledge resource to navigate through and explore. Biozon unifies multiple biological databases consisting of a variety of data types (such as DNA sequences, proteins, interactions and cellular pathways). It is fundamentally different from previous efforts as it uses a single extensive and tightly connected graph schema wrapped with hierarchical ontology of documents and relations. Beyond warehousing existing data, Biozon computes and stores novel derived data, such as similarity relationships and functional predictions. The integration of similarity data allows propagation of knowledge through inference and fuzzy searches. Sophisticated methods of query that span multiple data types were implemented and first-of-a-kind biological ranking systems were explored and integrated. CONCLUSION: The Biozon system is an extensive knowledge resource of heterogeneous biological data. Currently, it holds more than 100 million biological documents and 6.5 billion relations between them. The database is accessible through an advanced web interface that supports complex queries, "fuzzy" searches, data materialization and more, online at http://biozon.org.

Animals↗

PathwayVoyager: pathway mapping using the Kyoto Encyclopedia of Genes and Genomes (KEGG) database.

BACKGROUND: Equally important and challenging as genome annotation, is the subsequent classification of predicted genes into their respective pathways. The Kyoto Encyclopedia of Genes and Genomes (KEGG) represents a database consisting of known genes and their respective biochemical functionalities. Although accessible online, analyses of multiple genes are time consuming and are not suitable for analyzing data sets that are proprietary. RESULTS: Presented here is a new software solution that utilizes the KEGG online database for pathway mapping of partial and whole prokaryotic genomes. PathwayVoyager retrieves user-defined subsets of the KEGG database and stores the data as local, blast-formatted databases. Previously selected datasets can be re-used, reducing run-time significantly. Whole or partial genomes can be automatically analyzed using NCBI's BlastP algorithm and ORFs with similarities below the user-defined threshold will be marked on pathway maps. Multiple gene hits are sorted by similarity. Since no sequence information is transmitted over the Internet, PathwayVoyager is an ideal solution for pathway mapping and reconstruction of confidential DNA sequence data. CONCLUSION: PathwayVoyager represents an alternative approach to many already existing, more complex pathway reconstructions software solutions. This software does not require any dedicated hardware or software and is flexible and straightforward to use. It is ideally suited for environments where analyses on variable datasets are desired.

Computational Biology↗

AgBase: a functional genomics resource for agriculture.

BACKGROUND: Many agricultural species and their pathogens have sequenced genomes and more are in progress. Agricultural species provide food, fiber, xenotransplant tissues, biopharmaceuticals and biomedical models. Moreover, many agricultural microorganisms are human zoonoses. However, systems biology from functional genomics data is hindered in agricultural species because agricultural genome sequences have relatively poor structural and functional annotation and agricultural research communities are smaller with limited funding compared to many model organism communities. DESCRIPTION: To facilitate systems biology in these traditionally agricultural species we have established "AgBase", a curated, web-accessible, public resource http://www.agbase.msstate.edu for structural and functional annotation of agricultural genomes. The AgBase database includes a suite of computational tools to use GO annotations. We use standardized nomenclature following the Human Genome Organization Gene Nomenclature guidelines and are currently functionally annotating chicken, cow and sheep gene products using the Gene Ontology (GO). The computational tools we have developed accept and batch process data derived from different public databases (with different accession codes), return all existing GO annotations, provide a list of products without GO annotation, identify potential orthologs, model functional genomics data using GO and assist proteomics analysis of ESTs and EST assemblies. Our journal database helps prevent redundant manual GO curation. We encourage and publicly acknowledge GO annotations from researchers and provide a service for researchers interested in GO and analysis of functional genomics data. CONCLUSION: The AgBase database is the first database dedicated to functional genomics and systems biology analysis for agriculturally important species and their pathogens. We use experimental data to improve structural annotation of genomes and to functionally characterize gene products. AgBase is also directly relevant for researchers in fields as diverse as agricultural production, cancer biology, biopharmaceuticals, human health and evolutionary biology. Moreover, the experimental methods and bioinformatics tools we provide are widely applicable to many other species including model organisms.

Agriculture↗

A proposed architecture and method of operation for improving the protection of privacy and confidentiality in disease registers.

BACKGROUND: Disease registers aim to collect information about all instances of a disease or condition in a defined population of individuals. Traditionally methods of operating disease registers have required that notifications of cases be identified by unique identifiers such as social security number or national identification number, or by ensembles of non-unique identifying data items, such as name, sex and date of birth. However, growing concern over the privacy and confidentiality aspects of disease registers may hinder their future operation. Technical solutions to these legitimate concerns are needed. DISCUSSION: An alternative method of operation is proposed which involves splitting the personal identifiers from the medical details at the source of notification, and separately encrypting each part using asymmetrical (public key) cryptographic methods. The identifying information is sent to a single Population Register, and the medical details to the relevant disease register. The Population Register uses probabilistic record linkage to assign a unique personal identification (UPI) number to each person notified to it, although not necessarily everyone in the entire population. This UPI is shared only with a single trusted third party whose sole function is to translate between this UPI and separate series of personal identification numbers which are specific to each disease register. SUMMARY: The system proposed would significantly improve the protection of privacy and confidentiality, while still allowing the efficient linkage of records between disease registers, under the control and supervision of the trusted third party and independent ethics committees. The proposed architecture could accommodate genetic databases and tissue banks as well as a wide range of other health and social data collections. It is important that proposals such as this are subject to widespread scrutiny by information security experts, researchers and interested members of the general public, alike.

Computer Security↗

International Committee on Taxonomy of Viruses and the 3,142 unassigned species.

In 2005, ICTV (International Committee on Taxonomy of Viruses), the official body of the Virology Division of the International Union of Microbiological Societies responsible for naming and classifying viruses, will publish its latest report, the state of the art in virus nomenclature and taxonomy. The book lists more than 6,000 viruses classified in 1,950 species and in more than 391 different higher taxa. However, GenBank contains a staggering additional 3,142 "species" unaccounted for by the ICTV report. This paper reviews the reasons for such a situation and suggests what might be done in the near future to remedy this problem, particularly in light of the potential for a ten-fold increase in virus sequencing in the coming years that would generate many unclassified viruses. A number of changes could be made both at ICTV and GenBank to better handle virus taxonomy and classification in the future.

Advisory Committees↗

Annotation of the Drosophila melanogaster euchromatic genome: a systematic review.

BACKGROUND: The recent completion of the Drosophila melanogaster genomic sequence to high quality and the availability of a greatly expanded set of Drosophila cDNA sequences, aligning to 78% of the predicted euchromatic genes, afforded FlyBase the opportunity to significantly improve genomic annotations. We made the annotation process more rigorous by inspecting each gene visually, utilizing a comprehensive set of curation rules, requiring traceable evidence for each gene model, and comparing each predicted peptide to SWISS-PROT and TrEMBL sequences. RESULTS: Although the number of predicted protein-coding genes in Drosophila remains essentially unchanged, the revised annotation significantly improves gene models, resulting in structural changes to 85% of the transcripts and 45% of the predicted proteins. We annotated transposable elements and non-protein-coding RNAs as new features, and extended the annotation of untranslated (UTR) sequences and alternative transcripts to include more than 70% and 20% of genes, respectively. Finally, cDNA sequence provided evidence for dicistronic transcripts, neighboring genes with overlapping UTRs on the same DNA sequence strand, alternatively spliced genes that encode distinct, non-overlapping peptides, and numerous nested genes. CONCLUSIONS: Identification of so many unusual gene models not only suggests that some mechanisms for gene regulation are more prevalent than previously believed, but also underscores the complex challenges of eukaryotic gene prediction. At present, experimental data and human curation remain essential to generate high-quality genome annotations.

Animals↗

The past, present and future of genome-wide re-annotation.

Annotation, the process by which structural or functional information is inferred for genes or proteins, is crucial for obtaining value from genome sequences. We define the process of annotating a previously annotated genome sequence as 're-annotation', and examine the strengths and weaknesses of current manual and automatic genome-wide re-annotation approaches.

Computational Biology↗

REvolutionH-tl 2.0: A fast and robust tool for decoding evolutionary gene histories.

REvolutionH-tl is a fast, scalable, and integrated software platform for inferring orthology relationships, gene trees, species trees, and reconciled evolutionary scenarios directly from sequence data. Built upon the formal framework of best match graphs (BMGs), REvolutionH-tl predicts orthogroups and orthologous gene pairs with high accuracy, requiring neither precomputed trees nor multiple external tools. The software reconstructs event-labeled gene and species trees, seamlessly integrating reconciliation to produce fast, accurate, and biologically insightful evolutionary scenarios. Through extensive benchmarking on synthetic datasets with known ground truth, REvolutionH-tl outperforms or matches the accuracy of established tools such as OrthoFinder, Proteinortho, RAxML, GeneRax, and RANGER-DTL, while achieving significantly lower runtimes. A key innovation of REvolutionH-tl is its built-in support for detailed, publication-ready visualizations, which allow users to explore genome evolution dynamics, orthogroup composition, and reconciliation results with clarity and ease. These visual features position REvolutionH-tl as the first platform of its kind to combine analytical precision with intuitive interpretability. The software is open-source, cross-platform, and freely available at https://pypi.org/project/revolutionhtl/, providing a robust solution for large-scale evolutionary analyses in comparative genomics.

Software↗

MPID-T: database for sequence-structure-function information on T-cell receptor/peptide/MHC interactions.

UNLABELLED: Normal adaptive immune responses operate under major histocompatibility complex (MHC) restriction by binding to specific, short antigenic peptides and presenting them to appropriate T-cell receptors (TcRs). Sequence-structure-function information is critical in understanding the principles governing peptide/MHC (pMHC) and TcR/pMHC recognition and binding. A new database for sequence-structure-function information on TcR/pMHC interactions, MHC-Peptide Interaction Database version T (MPID-T), is now available with the latest available Protein Data Bank (PDB) data and interaction parameters on TcR/pMHC complexes. MPID-T is a manually curated MySQL database containing experimentally determined structures of 187 pMHC complexes and 16 TcR/pMHC complexes available in the PDB. Each structure is manually verified, classified, and analysed for intermolecular interactions (i) between the MHC and its corresponding bound peptide and (ii) between TcR and its bound pMHC complex where TcR structural information is available. The MPID-T database retrieval system has precomputed interaction parameters that include solvent accessibility, hydrogen bonds, gap volume and gap index. Structural visualisation of the TcR/pMHC complex, pMHC complex, MHC or the bound peptide can be performed using freely available graphics applications such as MDL Chime or RasMol, while structural alignment (based on MHC class and peptide length) can be viewed using the Jmol molecular viewer or an MDL Chime-compatible web browser client. MPID-T contains structural descriptors for in-depth characterisation of TcR/pMHC and pMHC interactions. The ultimate purpose of MPID-T is to enhance the understanding of the binding mechanism underlying TcR/pMHC and pMHC interactions by mapping the TcR footprint on the MHC and its bound peptide, as this eventually determines T-cell recognition and binding. AVAILABILITY: The MPID-T database retrieval system is available at http://surya.bic.nus.edu.sg/mpidt CONTACT: Joo Chuan Tong (jctong@i2r.a-star.edu.sg).

Animals↗

Can ends justify the means? Digging deep for human fusion genes of prokaryotic origin.

Gene fusion has been described as an important evolutionary phenomenon. This report focuses on identifying, analyzing, and tabulating human fusion proteins of prokaryotic origin. These fusion proteins are found to mimic operons, simulate protein-protein interfaces in prokaryotes, exhibiting multiple functions and alternative splicing in humans. The accredited biological functions for each of these proteins is made available as a database at http://sege.ntu.edu.sg/wester/fusion/

Alternative Splicing↗

Instabilotyping: comprehensive identification of frameshift mutations caused by coding region microsatellite instability.

Coding region frameshift mutation caused by microsatellite instability (MSI) is one mechanism contributing to tumorigenesis in cancers with MSI in high frequency. Mutation of TGFBR2 is one example of this process. To identify additional examples, a large-scale genomic screen of coding region microsatellites was conducted. 1115 coding homopolymeric loci with six or more nucleotides were identified in an online genetic database. Mutational screening was performed at 152 of these loci in 46 colorectal tumors with MSI in high frequency. Nine loci were mutated in > or =20% of tumors, 10 loci in 10-20%, 24 loci in 5-10%, 43 loci in <5%, and 66 loci were not mutated in any tumors. The most frequently mutated novel loci were the activin type II receptor gene (58.1%), SEC63 (48.8%), AIM 2 (47.6%), a gene encoding a subunit of the NADH-ubiquinone oxidoreductase complex (27.9%), a homologue of mouse cordon-bleu (23.8%), and EBP1/PA2G4 (20.9%). This genome-wide approach identifies coding region MSI in genes or pathways not implicated previously in colorectal tumorigenesis, which may merit functional study or other additional analysis.

3' Untranslated Regions↗

Transposon tagging and the study of root development in Arabidopsis.

The maize Ac-Ds transposable element family has been used as the basis of transposon mutagenesis systems that function in a variety of plants, including Arabidopsis. We have developed modified transposons and methods which simplify the detection, cloning and analysis of insertion mutations. We have identified and are analyzing two plant lines in which genes expressed either in the root cap cells or in the quiescent cells, cortex/endodermal initial cells and columella cells of the root cap have been tagged with a transposon carrying a reporter gene. A gene expressed in root cap cells tagged with an enhancer-trap Ds was isolated and its corresponding EST cDNA was identified. Nucleotide and deduced amino acid sequences of the gene show no significant similarity to other genes in the database. Genetic ablation experiments have been done by fusing a root cap-specific promoter to the diphtheria toxin A-chain gene and introducing the fusion construct into Arabidopsis plants. We find that in addition to eliminating gravitropism, root cap ablation inhibits elongation of roots by lowering root meristematic activities.

Arabidopsis↗

Large-scale open bioinformatics data resources.

The data explosion in bioinformatics is relentless. More and more genomes are being sequenced and many new types of datasets are being generated in large-scale projects. Integration and true open access to the data are still difficult issues, although they are gradually being addressed. Notably, certain fields have good standardization and interoperability, while others lag behind. This review summarizes the latest developments in genome and sequences databases, transcriptomics data (ESTs, ORESTES, full-length cDNAs), proteomics data (protein databases, protein structures, family and domain classification) as well as loosely integrated fields, such as microarray experiments, mutation databases and databases of regulatory regions and elements. The review attempts to resist simply summarizing what data are available, and aims to provide a critical look at some of the integration and access issues associated with several of these resources.

Computational Biology↗