Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Databases, Genetic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

Structural and functional characterization of gene products encoded in the human genome by homology detection.

Availability of the human genome data has enabled the exploration of a huge amount of biological information encoded in it. There are extensive ongoing experimental efforts to understand the biological functions of the gene products encoded in the human genome. However, computational analysis can aid immensely in the interpretation of biological function by associating known functional/structural domains to the human proteins. In this article we have discussed the implications of such associations. The association of structural domains to human proteins could help in prioritizing the targets for structure determination in the structural genomics initiatives. The protein kinase family is one of the most frequently occurring protein domain families in the human proteome while P-loop hydrolase, which comprises many GTPases and ATPases, is a highly represented superfamily. Using the superfamily relationships between families of unknown and known structures we could increase structural information content of the human genome by about 5%. We could also make new associations of domain families to 33 human proteins that are potentially linked to genetically inherited diseases.

Databases, Genetic↗

BacS: an abundant bacteroid protein in Rhizobium etli whose expression ex planta requires nifA.

Rhizobium etli CFN42 bacteroids from bean nodules possessed an abundant 16-kDa protein (BacS) that was found in the membrane pellet after cell disruption. This protein was not detected in bacteria cultured in tryptone-yeast extract. In minimal media, it was produced at low oxygen concentration but not in a mutant whose nifA was disrupted. N-terminal sequencing of the protein led to isolation of a bacS DNA fragment. DNA hybridization and nucleotide sequencing revealed three copies of the bacS gene, all residing on the main symbiotic plasmid of strain CFN42. A stretch of 304 nucleotides, exactly conserved upstream of all three bacS open reading frames, had very close matches with the NifA and sigma 54 consensus binding sequences. The only bacS homology in the genetic sequence databases was to three hypothetical proteins of unknown function, all from rhizobial species. Mutation and genetic complementation indicated that each of the bacS genes gives rise to a BacS polypeptide. Mutants disrupted or deleted in all three genes did not produce the BacS polypeptide but were Nod+ and Fix+ on Phaseolus vulgaris.

Aerobiosis↗

Mining OMIM for insight into complex diseases.

Understanding clinical phenotypes through their corresponding genotypes is one of the principal goals of genetic research. Though achieving this goal is relatively simple with single gene syndromes, more complex diseases often consist of varied clinical phenotypes that may be the result of interactions among multiple genetic loci. Microarray technology has brought the phenotype -genotype relationship to the molecular level, using differently behaving cancers, for example, as the basis for comparing patterns of gene expression. With this feasibility study, we attempted to use similar methods of analysis at the clinical level, in order to evaluate our hypothesis that the clustering of clinical phenotypes would provide information that would be useful in elucidating their underlying genotypes. Because of its breadth of content and detailed descriptions, we used OMIM as our source material for phenotypic and genetic information. After processing the source material, we then performed self-organizing map and hierarchical clustering analysis on representative diseases by phenotypic category. Through pre-determined queries over this analysis, we made two findings of potential clinical significance, one concerning diabetes and another concerning progressive neurologic diseases. Our methods provide a formal approach to analyzing phenotypes among diverse diseases, and may help indicate fruitful areas for further research into their underlying genetic causes.

Cluster Analysis↗

Constructing epigenetic regulatory landscapes of plant lncRNAs-an exploration utilizing the novel specialized platform PERlncDB.

Long non-coding RNAs (lncRNAs), once overlooked as transcriptional byproducts, are now recognized for their crucial roles in plant growth, development, and stress responses, with increasing focus on their epigenetic regulation. However, studies investigating epigenomic signals to explore the functions of lncRNAs in plants remain relatively limited. This study collected a comprehensive dataset of over 160 000 high-quality lncRNAs from 19 representative plant species and integrated 6715 ChIP-seq, BS-seq, and RNA-seq datasets to analyze epigenomic patterns at lncRNA loci. Results showed elevated DNA methylation in lncRNA regions. The highest levels occurred in transposable element-associated lncRNAs. Additionally, activating histone modifications at lncRNA loci showed tissue specificity, with epigenetic preferences differed from those at protein-coding gene (PCG) loci. Differential site analysis in epigenetic mutants further highlighted the selective regulation of lncRNA loci by specific epigenetic factors. To facilitate research, we developed PERlncDB, a platform that provides species-specific lncRNA browsing, epigenetic annotation, cross-species conservation analysis, and visualization of epigenomic landscapes. Case studies on MARS and LINC-AP2 emphasized the platform's utility. Conserved epigenetic mechanisms regulating lncRNAs across species, exemplified by a syntenic conserved MET1-regulated lncRNA pair in Arabidopsis and tomato, suggested the stability of regulatory mechanisms underlying lncRNA functions. This work provides critical insights and resources for understanding plant lncRNA epigenetic regulation.

RNA, Long Noncoding↗

Identification of genes associated with natural competence in Helicobacter pylori by transposon shuttle random mutagenesis.

To identify genes involved in DNA transformation, we generated 1500 insertion mutants of a Helicobacter pylori strain by transposon shuttle mutagenesis. All mutant strains were screened for their frequency of natural transformation. A total of 20 mutant strains were found to exhibit a significantly decreased transformation frequency. DNA sequencing revealed seven genetic loci, including the reported comB locus, HP0017 (a putative virB4 homologue) and five loci without database match (HP0015, HP1089, HP1326, HP1424, and HP1473) from the 20 mutants. Reknockout of HP1326 revealed no impairment in natural transformation, while the other 5 mutants showed the same defective in natural transformation. Mutation of HP0017 severely impaired natural transformation both chromosome and plasmid DNA. Slot blot analysis revealed that some noncompetent strains had decreased virB4 RNA expression levels compared with competent strains. Nineteen ORFs had decreased expression levels in virB4 knockout mutant by microarray. Therefore, our data indicate that HP0017 is a virB4 homologue and is essential in the natural competence of H. pylori. HP0015, HP1089, HP1424, and HP1473 genes could be also involved in natural transformation.

Blotting, Southern↗

Up-to-date, and taxonomy-curated mcrA reference databases for methanogen community profiling.

The methyl-coenzyme M reductase subunit alpha gene (mcrA) is an important phylogenetic marker for high throughput ecological profiling of methanogenic archaea, central to industrial biological methane production and greenhouse gas emissions. Yet, dedicated reference databases predate current relevant NCBI sequence accumulation and archaeal taxonomic revision. We present three updated mcrA reference databases: (i) one derived from NCBI-catalogued methanogen genomes (1572 sequences); (ii) a database built by expansion of a previously published reference dataset, leveraging the NCBI nucleotide collection (27,942 sequences); (iii) a curated-taxonomy version of the latter. The updated amplicon databases provide a ∼ 3.5-fold sequence richness expansion, extend genus-level richness from 31 to 83 taxa, more than 4-fold species-level richness, and incorporate novel lineages compared with the previous reference dataset (e.g. Thermoplasmatota-encompassed). All databases were formatted to support analysis with relevant contemporary software pipelines and packages. Overall, the generated databases facilitate a highly improved characterization of methanogen diversity and ecology.

Archaea↗

AutoPM3: enhancing variant interpretation via LLM-driven PM3 evidence extraction from scientific literature.

MOTIVATION: Rare diseases affect over 300 million people worldwide and are often caused by genetic variants. While variant detection has become cost-effective, interpreting these variants-particularly collecting literature-based evidence like ACMG/AMP PM3-remains complex and time-consuming. RESULTS: We present AutoPM3, a method that automates PM3 evidence extraction from literatures using open-source large language models (LLMs). AutoPM3 combines a Text2SQL-based variant extractor and a retrieval-augmented generation (RAG) module, enhanced by a variant-specific retriever and fine-tuned LLM, to separately process tables and text. We curated PM3-Bench, a dataset of 1027 variant-publication evidence pairs from ClinGen. On openly accessible pairs, AutoPM3 achieved 86.1% accuracy for variant hits and 72.5% recall for in trans variants-outperforming other methods, including those using larger models. We uncovered the effectiveness of AutoPM3's key modules, especially for variant-specific retriever and Text2SQL, through the sequential ablation study. AutoPM3 located evidence in 76 s, demonstrating that open-source LLMs can offer an efficient, cost-effective solution for rare disease diagnosis. AVAILABILITY AND IMPLEMENTATION: AutoPM3 is implemented and freely available under the MIT license at https://github.com/HKU-BAL/AutoPM3.

Genetic Variation↗

Comparative genome analysis of the yellow fever mosquito Aedes aegypti with Drosophila melanogaster and the malaria vector mosquito Anopheles gambiae.

An in silico comparative genomics approach was used to identify putative orthologs to genetically mapped genes from the mosquito, Aedes aegypti, in the Drosophila melanogaster and Anopheles gambiae genome databases. Comparative chromosome positions of 73 D. melanogaster orthologs indicated significant deviations from a random distribution across each of the five A. aegypti chromosomal regions, suggesting that some ancestral chromosome elements have been conserved. However, the two genomes also reflect extensive reshuffling within and between chromosomal regions. Comparative chromosome positions of A. gambiae orthologs indicate unequivocally that A. aegypti chromosome regions share extensive homology to the five A. gambiae chromosome arms. Whole-arm or near-whole-arm homology was contradicted with only two genes among the 75 A. aegypti genes for which orthologs to A. gambiae were identified. The two genomes contain large conserved chromosome segments that generally correspond to break/fusion events and a reciprocal translocation with extensive paracentric inversions evident within. Only very tightly linked genes are likely to retain conserved linear orders within chromosome segments. The D. melanogaster and A. gambiae genome databases therefore offer limited potential for comparative positional gene determinations among even closely related dipterans, indicating the necessity for additional genome sequencing projects with other dipteran species.

Aedes↗

Composite genome map and recombination parameters derived from three archetypal lineages of Toxoplasma gondii.

Toxoplasma gondii is a highly successful protozoan parasite in the phylum Apicomplexa, which contains numerous animal and human pathogens. T.gondii is amenable to cellular, biochemical, molecular and genetic studies, making it a model for the biology of this important group of parasites. To facilitate forward genetic analysis, we have developed a high-resolution genetic linkage map for T.gondii. The genetic map was used to assemble the scaffolds from a 10X shotgun whole genome sequence, thus defining 14 chromosomes with markers spaced at approximately 300 kb intervals across the genome. Fourteen chromosomes were identified comprising a total genetic size of approximately 592 cM and an average map unit of approximately 104 kb/cM. Analysis of the genetic parameters in T.gondii revealed a high frequency of closely adjacent, apparent double crossover events that may represent gene conversions. In addition, we detected large regions of genetic homogeneity among the archetypal clonal lineages, reflecting the relatively few genetic outbreeding events that have occurred since their recent origin. Despite these unusual features, linkage analysis proved to be effective in mapping the loci determining several drug resistances. The resulting genome map provides a framework for analysis of complex traits such as virulence and transmission, and for comparative population genetic studies.

Animals↗

Efficient gene-driven germ-line point mutagenesis of C57BL/6J mice.

BACKGROUND: Analysis of an allelic series of point mutations in a gene, generated by N-ethyl-N-nitrosourea (ENU) mutagenesis, is a valuable method for discovering the full scope of its biological function. Here we present an efficient gene-driven approach for identifying ENU-induced point mutations in any gene in C57BL/6J mice. The advantage of such an approach is that it allows one to select any gene of interest in the mouse genome and to go directly from DNA sequence to mutant mice. RESULTS: We produced the Cryopreserved Mutant Mouse Bank (CMMB), which is an archive of DNA, cDNA, tissues, and sperm from 4,000 G1 male offspring of ENU-treated C57BL/6J males mated to untreated C57BL/6J females. Each mouse in the CMMB carries a large number of random heterozygous point mutations throughout the genome. High-throughput Temperature Gradient Capillary Electrophoresis (TGCE) was employed to perform a 32-Mbp sequence-driven screen for mutations in 38 PCR amplicons from 11 genes in DNA and/or cDNA from the CMMB mice. DNA sequence analysis of heteroduplex-forming amplicons identified by TGCE revealed 22 mutations in 10 genes for an overall mutation frequency of 1 in 1.45 Mbp. All 22 mutations are single base pair substitutions, and nine of them (41%) result in nonconservative amino acid substitutions. Intracytoplasmic sperm injection (ICSI) of cryopreserved spermatozoa into B6D2F1 or C57BL/6J ova was used to recover mutant mice for nine of the mutations to date. CONCLUSIONS: The inbred C57BL/6J CMMB, together with TGCE mutation screening and ICSI for the recovery of mutant mice, represents a valuable gene-driven approach for the functional annotation of the mammalian genome and for the generation of mouse models of human genetic diseases. The ability of ENU to induce mutations that cause various types of changes in proteins will provide additional insights into the functions of mammalian proteins that may not be detectable by knockout mutations.

Animals↗

An inquiry into protein structure and genetic disease: introducing undergraduates to bioinformatics in a large introductory course.

This inquiry-based lab is designed around genetic diseases with a focus on protein structure and function. To allow students to work on their own investigatory projects, 10 projects on 10 different proteins were developed. Students are grouped in sections of 20 and work in pairs on each of the projects. To begin their investigation, students are given a cDNA sequence that translates into a human protein with a single mutation. Each case results in a genetic disease that has been studied and recorded in the Online Mendelian Inheritance in Man (OMIM) database. Students use bioinformatics tools to investigate their proteins and form a hypothesis for the effect of the mutation on protein function. They are also asked to predict the impact of the mutation on human physiology and present their findings in the form of an oral report. Over five laboratory sessions, students use tools on the National Center for Biotechnology Information (NCBI) Web site (BLAST, LocusLink, OMIM, GenBank, and PubMed) as well as ExPasy, Protein Data Bank, ClustalW, the Kyoto Encyclopedia of Genes and Genomes (KEGG) database, and the structure-viewing program DeepView. Assessment results showed that students gained an understanding of the Web-based databases and tools and enjoyed the investigatory nature of the lab.

Algorithms↗

Genome-wide linkage disequilibrium and haplotype maps.

There is currently a broad effort to produce genome-wide high-density linkage disequilibrium (LD) maps with single nucleotide polymorphisms. The hope is that the resulting maps can be exploited to find genes that affect the onset and severity of at least some common human diseases. These maps may also be useful for identifying genes that affect drug response or the likelihood of drug toxicities. The goal of this review is to provide a broad overview of some of the key concerns motivating the design of a major international project called the International Haplotype Map Project. The process of map production requires the identification of very large numbers of polymorphic sites, implementation of facile, highly accurate and inexpensive genotyping production pipelines, and provision for public access to the genotype data. Great progress has been made recently in genotyping methods and these advances are allowing very large-scale data collection. A major goal of these efforts is to enable the selection of subsets of markers that capture useful genetic information in short genomic intervals, while optimally reducing the number of markers that must be genotyped. Standard measures of LD provide a starting point but may not fully capture the complexity of the information inherent in the data. Extremely dense genotype data in several broadly representative populations (European, Chinese, Japanese, and Yoruba) should yield important insights into the genetic structure of most genes. Further study is required to determine how broadly applicable the data will be to other population groups. Significant challenges lie ahead in determining the best methods for the selection of markers in disease/phenotype studies, large-scale genotyping, and analysis of the resulting genetic data.

Animals↗

Mapping the proteome of barrel medic (Medicago truncatula).

A survey of six organ-/tissue-specific proteomes of the model legume barrel medic (Medicago truncatula) was performed. Two-dimensional polyacrylamide gel electrophoresis reference maps of protein extracts from leaves, stems, roots, flowers, seed pods, and cell suspension cultures were obtained. Five hundred fifty-one proteins were excised and 304 proteins identified using peptide mass fingerprinting and matrix-assisted laser desorption ionization time-of-flight mass spectrometry. Nanoscale high-performance liquid chromatography coupled with tandem quadrupole time-of-flight mass spectrometry was used to validate marginal matrix-assisted laser desorption ionization time-of-flight mass spectrometry protein identifications. This dataset represents one of the most comprehensive plant proteome projects to date and provides a basis for future proteome comparison of genetic mutants, biotically and abiotically challenged plants, and/or environmentally challenged plants. Technical details concerning peptide mass fingerprinting, database queries, and protein identification success rates in the absence of a sequenced genome are reported and discussed. A summary of the identified proteins and their putative functions are presented. The tissue-specific expression of proteins and the levels of identified proteins are compared with their related transcript abundance as quantified through EST counting. It is estimated that approximately 50% of the proteins appear to be correlated with their corresponding mRNA levels.

Amino Acid Sequence↗

[The informatics of human genome and traditional Chinese medicine].

Guided by the theory and methodology of yin-yang set derived from Changing Book and Medicine Canon, and using genetics as a bridge, we have tried to bring together the ancient functional systematology and modern structural one as well as Eastern and Western medicine, thereby promoting the modernization of traditional Chinese medicine (TCM) in theory and in clinical practice. Herein, we used virtual technology to transform the genetic information in OMIM of NCBI (National Center for Biotechnology Information of USA, http://www.ncbi.nlm.nih.gov ) into a secondary database in the form of webpages. There are sixteen kinds of the database named gene morbidity ones as followings as: the nature of gene, the profile of common phenotype, a interaction of endogenous, the disease of a organ or a viscera pathogenesis phenomenon, TCM, the sign of diagnosis of western medicine, the gene response to environment, syndrome, disease, nerve and -endocrine, tumor and cancer, psychology and behavior, morbidity, endo-factor of molecular information, expression, the interaction between endogenous and exogenous in which there is 4 711 words, files. The advantages of the database are its aptness for using human fuzzy intelligence to recognize things, suitability to uncovering the noumenon (yinyang) nature of an object and applicability to clinical use.

Computational Biology↗