Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Databases, Genetic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40Linked to original sources

A combined linkage-physical map of the human genome.

We have constructed de novo a high-resolution genetic map that includes the largest set, to our knowledge, of polymorphic markers (N=14,759) for which genotype data are publicly available; that combines genotype data from both the Centre d'Etude du Polymorphisme Humain (CEPH) and deCODE pedigrees; that incorporates single-nucleotide polymorphisms; and that also incorporates sequence-based positional information. The position of all markers on our map is corroborated by both genomic sequence and recombination-based data. This specific combination of features maximizes marker inclusion, coverage, and resolution, making this map uniquely suitable as a comprehensive resource for determining genetic map information (order and distances) for any large set of polymorphic markers.

Chromosome Mapping↗

Evaluating outlier loci and their effect on the identification of pedigree errors.

Homozygosity outlier loci, which show patterns of variation that are extremely divergent from the rest of the genome, can be evaluated by comparison of the homozygosity under Hardy-Weinberg proportions (the sum of the squares of allele frequencies) with the expected homozygosity under neutrality. Such outlier loci are potentially under selection (balancing selection or directional selection) when genome-wide effects (such as bottleneck and rapid population growth) are excluded. Outlier loci show skewed allele frequencies with respect to neutrality and may therefore affect the identification of pedigree errors. However, choosing neutral markers (excluding outlier loci) for the identification of pedigree errors has been neglected thus far. Our results showed that 4.1%, 5.5%, and 1.5% of the microsatellite markers, Illumina single-nucleotide polymorphisms (SNPs), and Affymetrix SNPs, respectively, on the autosomes appear to be under balancing selection (p or=40%) appear to be under balancing selection. Pedigree structure errors in 15 of 143 pedigrees were detected using microsatellite markers from the autosomes and/or selected SNPs from chromosomes 1 to 18 of the Illumina and/or selected SNPs from chromosomes 1 to 16 of the Affymetrix. Outlier loci did not make a major difference to the identification of pedigree errors. The Collaborative Study on the Genetics of Alcoholism data has pedigree errors and some of them may be due to sample mix up.

Alcoholism↗

Gene selection based on multi-class support vector machines and genetic algorithms.

Microarrays are a new technology that allows biologists to better understand the interactions between diverse pathologic state at the gene level. However, the amount of data generated by these tools becomes problematic, even though data are supposed to be automatically analyzed (e.g., for diagnostic purposes). The issue becomes more complex when the expression data involve multiple states. We present a novel approach to the gene selection problem in multi-class gene expression-based cancer classification, which combines support vector machines and genetic algorithms. This new method is able to select small subsets and still improve the classification accuracy.

Algorithms↗

MEGA3: Integrated software for Molecular Evolutionary Genetics Analysis and sequence alignment.

With its theoretical basis firmly established in molecular evolutionary and population genetics, the comparative DNA and protein sequence analysis plays a central role in reconstructing the evolutionary histories of species and multigene families, estimating rates of molecular evolution, and inferring the nature and extent of selective forces shaping the evolution of genes and genomes. The scope of these investigations has now expanded greatly owing to the development of high-throughput sequencing techniques and novel statistical and computational methods. These methods require easy-to-use computer programs. One such effort has been to produce Molecular Evolutionary Genetics Analysis (MEGA) software, with its focus on facilitating the exploration and analysis of the DNA and protein sequence variation from an evolutionary perspective. Currently in its third major release, MEGA3 contains facilities for automatic and manual sequence alignment, web-based mining of databases, inference of the phylogenetic trees, estimation of evolutionary distances and testing evolutionary hypotheses. This paper provides an overview of the statistical methods, computational tools, and visual exploration modules for data input and the results obtainable in MEGA.

Databases, Genetic↗

Toward a comprehensive set of asthma susceptibility genes.

Epidemiological and twin studies have demonstrated that asthma is under genetic and environmental influences. Numerous candidate gene association studies as well as genome-wide linkage scans have followed, aiming to elucidate the genetic architecture underlying this complex disease. Several promising asthma susceptibility genes were identified, and a comprehensive catalogue of these genes seems a realistic goal within 5 to 10 years. However, a key challenge is to understand the combination of genes and environmental factors that gives rise to the disease in a specific individual. Currently, most of the reports of asthma susceptibility genes are either preliminary or controversial, with little knowledge about the genetic mechanisms leading to abnormal function of the gene that promotes the development of asthma. Replications of published associations are relatively few. Many factors, including the inherent complexity of asthma as well as methodological issues, can explain these inconsistencies. Promising genetic tools are emerging with the completion of the International HapMap Project that will increase the scope of gene-discovery investigations. It is hoped that these tools, combined with validation studies in additional populations, will enable the creation of a comprehensive catalogue of susceptibility genes for asthma. Notwithstanding the difficulties in making sense of the vast amount of new genetic data, we already see the emergence of new biological pathways of atopy, airway remodeling, and asthma that may lead to novel therapeutic approaches.

Asthma↗

B.E.A.R. GeneInfo: a tool for identifying gene-related biomedical publications through user modifiable queries.

BACKGROUND: Once specific genes are identified through high throughput genomics technologies there is a need to sort the final gene list to a manageable size for validation studies. The triaging and sorting of genes often relies on the use of supplemental information related to gene structure, metabolic pathways, and chromosomal location. Yet in disease states where the genes may not have identifiable structural elements, poorly defined metabolic pathways, or limited chromosomal data, flexible systems for obtaining additional data are necessary. In these situations having a tool for searching the biomedical literature using the list of identified genes while simultaneously defining additional search terms would be useful. RESULTS: We have built a tool, BEAR GeneInfo, that allows flexible searches based on the investigators knowledge of the biological process, thus allowing for data mining that is specific to the scientist's strengths and interests. This tool allows a user to upload a series of GenBank accession numbers, Unigene Ids, Locuslink Ids, or gene names. BEAR GeneInfo takes these IDs and identifies the associated gene names, and uses the lists of gene names to query PubMed. The investigator can add additional modifying search terms to the query. The subsequent output provides a list of publications, along with the associated reference hyperlinks, for reviewing the identified articles for relevance and interest. An example of the use of this tool in the study of human prostate cancer cells treated with Selenium is presented. CONCLUSIONS: This tool can be used to further define a list of genes that have been identified through genomic or genetic studies. Through the use of targeted searches with additional search terms the investigator can limit the list to genes that match their specific research interests or needs. The tool is freely available on the web at http://prostategenomics.org1, and the authors will provide scripts and database components if requested mdatta@mcw.edu

Databases, Genetic↗

Comparison of linkage and association strategies for quantitative traits using the COGA dataset.

Genome scans using dense single-nucleotide polymorphism (SNP) data have recently become a reality. It is thought that the increase in information content for linkage analysis as a result of the denser scans will help refine previously identified linkage regions and possibly identify new regions not identifiable using the sparser, microsatellite scans. In the context of the dense SNP scans, it is also possible to consider association strategies to provide even more information about potential regions of interest. To circumvent the multiple-testing issues inherent in association analysis, we use a recently developed strategy, implemented in PBAT, which screens the data to identify the optimal SNPs for testing, without biasing the nominal significance level. We compare the results from the PBAT analysis to that of quantitative linkage analysis on chromosome 4 using the Collaborative Study on the Genetics of Alcoholism data, as released through Genetic Analysis Workshop 14.

Alcoholism↗

Molecular markers: tools to improve genebank efficiency.

Possibilities for using molecular markers to improve genebank efficiency are increasingly present thanks to developments in genebanks and developments in molecular genetics. These possibilities relate to all aspects of genebank management: acquisition, maintenance, characterisation and utilisation. However, two pitfalls should be avoided. The first lies in the neutrality of the most generally used markers, making them less suitable for optimising genetic diversity. The second is related to the considerable costs involved in using molecular markers. In many cases an economical analysis will have to decide if the markers can routinely be used in genebank operations. Some examples of model studies and applications of molecular markers in genebank operations will be presented, in which both genetic and economic aspects will be illustrated briefly. These examples involved existing genebank collections of wild lettuce, cabbage and wild potato.

Brassica↗

Bayesian approach to discovering pathogenic SNPs in conserved protein domains.

The success rate of association studies can be improved by selecting better genetic markers for genotyping or by providing better leads for identifying pathogenic single nucleotide polymorphisms (SNPs) in the regions of linkage disequilibrium with positive disease associations. We have developed a novel algorithm to predict pathogenic single amino acid changes, either nonsynonymous SNPs (nsSNPs) or missense mutations, in conserved protein domains. Using a Bayesian framework, we found that the probability of a microbial missense mutation causing a significant change in phenotype depended on how much difference it made in several phylogenetic, biochemical, and structural features related to the single amino acid substitution. We tested our model on pathogenic allelic variants (missense mutations or nsSNPs) included in OMIM, and on the other nsSNPs in the same genes (from dbSNP) as the nonpathogenic variants. As a result, our model predicted pathogenic variants with a 10% false-positive rate. The high specificity of our prediction algorithm should make it valuable in genetic association studies aimed at identifying pathogenic SNPs.

Algorithms↗

Information overload: assigning genetic functionality in the age of genomics and large-scale screening.

As more and more genome sequences are completed, it is becoming increasingly evident that our understanding of the function of most bacterial gene products is lacking. This is frustrating, particularly in the study of pathogens, where an understanding of the role of individual gene products would probably facilitate the development of novel antimicrobials and vaccines. Recently, we devised a technique known as virulence-attenuated pool (VAP) screening to help assign genetic functionality to gene products that the pathogen Vibrio cholerae requires for colonization. This screen and potential new applications of the VAP technique are discussed here.

Bacteria↗

Fungal genetic resource centres and the genomic challenge.

Fungal research and education has for many years been supported by public service genetic resource centres, whose roles have been to maintain, preserve and supply living cultures to the research community. In the genomic era, genetic resource centres are perhaps more important than ever before. The cultures held, many of which are described and validated by expert biosystematists, are valuable resources for the future. There is a need to supply genomic and proteomic research programmes with fully characterised organisms, as usage of organisms from unreliable sources can prove disastrous, not least in economical terms. However, mycologists often require more than just the organisms, for example, their associated information is vital for bioinformatic applications and some researchers may only require genomic DNA from the organism rather than the organism per se. Genetic resource centres are continually adapting to meet the needs of their users and the wider mycological research community, this associated with OECD international initiatives should ensure they exist to support research for many years to come. This review considers the impact of such initiatives, the current roles of fungal genetic resource centres, the mechanisms used to preserve organisms in a stable manner and the range of resources that are offered for genomic research.

Computational Biology↗

Bioinformatic insights from metagenomics through visualization.

Cutting-edge biological and bioinformatics research seeks a systems perspective through the analysis of multiple types of high-throughput and other experimental data for the same sample. Systems-level analysis requires the integration and fusion of such data, typically through advanced statistics and mathematics. Visualization is a complementary computational approach that supports integration and analysis of complex data or its derivatives. We present a bioinformatics visualization prototype, Juxter, which depicts categorical information derived from or assigned to these diverse data for the purpose of comparing patterns across categorizations. The visualization allows users to easily discern correlated and anomalous patterns in the data. These patterns, which might not be detected automatically by algorithms, may reveal valuable information leading to insight and discovery. We describe the visualization and interaction capabilities and demonstrate its utility in a new field, metagenomics, which combines molecular biology and genetics to identify and characterize genetic material from multi-species microbial samples.

Algorithms↗

Peeling off the hidden genetic heterogeneities of cancers based on disease-relevant functional modules.

Discovering molecular heterogeneities in phenotypically defined disease is of critical importance both for understanding pathogenic mechanisms of complex diseases and for finding efficient treatments. Recently, it has been recognized that cellular phenotypes are determined by the concerted actions of many functionally related genes in modular fashions. The underlying modular mechanisms should help the understanding of hidden genetic heterogeneities of complex diseases. We defined a putative disease module to be the functional gene groups in terms of both biological process and cellular localization, which are significantly enriched with genes highly variably expressed across the disease samples. As a validation, we used two large cancer datasets to evaluate the ability of the modules for correctly partitioning samples. Then, we sought the subtypes of complex diffuse large B-cell lymphoma (DLBCL) using a public dataset. Finally, the clinical significance of the identified subtypes was verified by survival analysis. In two validation datasets, we achieved highly accurate partitions that best fit the clinical cancer phenotypes. Then, for the notoriously heterogeneous DLBCL, we demonstrated that two partitioned subtypes using an identified module ("cellular response to stress") had very different 5-year overall rates (65% vs. 14%) and were highly significantly (P < 0.007) correlated with the clinical survival rate. Finally, we built a multivariate Cox proportional-hazard prediction model that included 4 genes as risk predictors for survival over DLBCL. The proposed modular approach is a promising computational strategy for peeling off genetic heterogeneities and understanding the modular mechanisms of human diseases such as cancers.

Databases, Genetic↗

Expressed sequence tag analysis of adult human lens for the NEIBank Project: over 2000 non-redundant transcripts, novel genes and splice variants.

PURPOSE: To explore the expression profile of the human lens and to provide a resource for microarray studies, expressed sequence tag (EST) analysis has been performed on cDNA libraries from adult lenses. METHODS: A cDNA library was constructed from two adult (40 year old) human lenses. Over two thousand clones were sequenced from the unamplified, un-normalized library. The library was then normalized and a further 2200 sequences were obtained. All the data were analyzed using GRIST (GRouping and Identification of Sequence Tags), a procedure for gene identification and clustering. RESULTS: The lens library (by) contains a low percentage of non-mRNA contaminants and a high fraction (over 75%) of apparently full length cDNA clones. Approximately 2000 reads from the unamplified library yields 810 clusters, potentially representing individual genes expressed in the lens. After normalization, the content of crystallins and other abundant cDNAs is markedly reduced and a similar number of reads from this library (fs) yields 1455 unique groups of which only two thirds correspond to named genes in GenBank. Among the most abundant cDNAs is one for a novel gene related to glutamine synthetase, which was designated "lengsin" (LGS). Analyses of ESTs also reveal examples of alternative transcripts, including a major alternative splice form for the lens specific membrane protein MP19. Variant forms for other transcripts, including those encoding the apoptosis inhibitor Livin and the armadillo repeat protein ARVCF, are also described. CONCLUSIONS: The lens cDNA libraries are a resource for gene discovery, full length cDNAs for functional studies and microarrays. The discovery of an abundant, novel transcript, lengsin, and a major novel splice form of MP19 reflect the utility of unamplified libraries constructed from dissected tissue. Many novel transcripts and splice forms are represented, some of which may be candidates for genetic diseases.

Adult↗

[The practice and explorations on optimization instructional design of genetics].

The Genetics teaching adopts instructional design improvement which centers on students to improve their understanding build-ups towards knowledge; attaches importance to the integration of the information technology and the course; utilizes effectively the genetics information resources on the Internet to foster students' information literacy and the ability of active learning; improves the operations of research experiment with the aim to foster the comprehensive ability; high-lights the bilateral activities of teaching and learning to create a harmonious and interactive teaching environment. Genetics teaching fulfils students' main-body function in order to realize the transformation between knowledge and ability to nurture excellent biology educators for the 21st century.

Databases, Genetic↗

Testing the chromosomal speciation hypothesis for humans and chimpanzees.

Fixed differences of chromosomal rearrangements between isolated populations may promote speciation by preventing between-population gene flow upon secondary contact, either because hybrids suffer from lowered fitness or, more likely, because recombination is reduced in rearranged chromosomal regions. This chromosomal speciation hypothesis thus predicts more rapid genetic divergence on rearranged than on colinear chromosomes because the former are less porous to gene flow. A number of studies of fungi, plants, and animals, including limited genetic data of humans and chimpanzees, support the hypothesis. Here we reexamine the hypothesis for humans and chimpanzees with substantially more genomic data than were used previously. No difference is observed between rearranged and colinear chromosomes in the level of genomic DNA sequence divergence between species. The same is also true for protein sequences. When the gorilla is used as an outgroup, no acceleration in protein sequence evolution associated with chromosomal rearrangements is found. Furthermore, divergence in expression pattern between orthologous genes is not significantly different for rearranged and colinear chromosomes. These results, showing that chromosomal rearrangements did not affect the rate of genetic divergence between humans and chimpanzees, are expected if incipient species on the evolutionary lineages separating humans and chimpanzees did not hybridize.

Animals↗

The Environmental Genome Project: phase I and beyond.

Human illness is caused by many interrelated factors including aging, inherited genetic predispositions, and a variety of environmental exposures. There is increasing awareness of the role of genetics as a factor that can dramatically alter susceptibility to all disease, especially environmentally induced chronic disease, such as cancer, asthma, diabetes, cardiovascular disease, and neurodegenerative disorders. In some cases, a genetic factor influences disease susceptibility in a small fraction of the population because it occurs at a low frequency or involves a relatively low-incidence disease; however, in other cases, a genetic factor increases susceptibility in a large number of individuals and involves a disease that occurs at high incidence, creating a large public health burden.

Aryldialkylphosphatase↗