Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “SNPs”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Choosing SNPs using feature selection.

A major challenge for genomewide disease association studies is the high cost of genotyping large number of single nucleotide polymorphisms (SNPs). The correlations between SNPs, however, make it possible to select a parsimonious set of informative SNPs, known as "tagging" SNPs, able to capture most variation in a population. Considerable research interest has recently focused on the development of methods for finding such SNPs. In this paper, we present an efficient method for finding tagging SNPs. The method does not involve computation-intensive search for SNP subsets but discards redundant SNPs using a feature selection algorithm. In contrast to most existing methods, the method presented here does not limit itself to using only correlations between SNPs in local groups. By using correlations that occur across different chromosomal regions, the method can reduce the number of globally redundant SNPs. Experimental results show that the number of tagging SNPs selected by our method is smaller than by using block-based methods. Supplementary website: http://htsnp.stanford.edu/FSFS/.

Algorithms↗

Comparative informativeness for linkage of multiple SNPs and single microsatellites.

Single nucleotide polymorphisms (SNPs), or biallelic markers, are popular in genetic linkage studies due to their abundance in the genome, stability, and ease of scoring. We determined the 'information ratio' (IR) of closely spaced SNPs in simulated nuclear families and affected sib pairs (ASPs). (The IR is the ratio of actual average maximum lod score to the maximum lod score attainable if the marker were fully informative.) The nuclear families included parental information, whereas the ASPs did not. We analyzed these SNPs in two ways: (1) using multipoint analysis, and (2) treating the SNPs as 'composite markers' (i.e., haplotypes, as assigned by GENEHUNTER). (3) We also calculated the IR of a single microsatellite marker with multiple alleles and compared with the IR from the SNPs. For each set of input conditions, we simulated 1000 nuclear families, of 2, 3, 4, or 5 children each, as well as 1000 ASPs. We generated SNP marker data for strings of k = 1, 2, 3, 5, 7, and 10 SNP loci, with no recombination (theta = 0) and no linkage disequilibrium among the SNPs. The MAF (minor allele frequency) was either 0.5 or 0.25, and allele frequencies were the same for all k loci in any analysis. We also generated marker data for one single-locus microsatellite marker, with m = 3, 4, 5, 6, 7, and 9 equally frequent alleles. In all simulations, the disease was fully penetrant dominant, and there was no recombination or linkage disequilibrium among markers or between marker and disease. When multipoint analysis was used, we found that 5-7 closely spaced SNPs were usually enough to yield an IR of approximately 100%, for nuclear families of any size. However, for the ASPs, even 7-10 SNPs yielded an IR of only 70-80%. A microsatellite with 9 equally frequent alleles yielded about the same IR (86-88%) as a string of 4-5 SNPs, in nuclear families. SNPs analyzed as 'composite markers' analyses performed worse, due to the inherent ambiguity of SNP haplotyping.

Alleles↗

Identification of single-nucleotide polymorphisms (SNPs) of human N-acetyltransferase genes NAT1, NAT2, AANAT, ARD1 and L1CAM in the Japanese population.

By direct sequencing of regions of the human genome containing five genes belonging to the acetyltransferase family, arylamine N-acetyltransferase (NAT1), arylamine N-acetyltransferase (NAT2), arylalkylamine N-acetyltransferase (AANAT), L1 cell adhesion molecule (L1CAM), and the human homolog of Saccharomyces cerevisiae N-acetyltransferase ARD1, we identified 53 single-nucleotide polymorphisms (SNPs) and two insertion/ deletion polymorphisms in 48 healthy Japanese volunteers. NAT1 and NAT2 are so-called drug-metabolizing enzymes. In the NAT1 gene we found two SNPs and a 3-bp insertion/ deletion polymorphism that corresponded to the NAT1*3, *10, and *18A/*18B alleles reported in other populations. The frequencies of NAT1* alleles in our Japanese subjects were 52.6% for NAT1*4, 1.0% for NAT1*3, 40.6% for NAT1*10, 2.6% for NAT1*18A and 3.1% for NAT1*18B. In the NAT2 gene we found 32 SNPs and a 1-bp insertion/ deletion polymorphism; 6 SNPs within the coding region were reported previously and belonged to the slow acetylator group (NAT2*5, NAT2*6 and NAT2*7), and 2 of the 8 SNPs in the 5' flanking region were reported in the dbSNP of GenBank, but the remaining 24 SNPs and the insertion/deletion polymorphism were novel. The frequencies of NAT2* alleles in Japanese (51.3% for NAT2*4, 1.6% for *5B, 26.1% for *6A, 2.2% for *6B, 1.2% for *7A, 10.1% for *7B, 7.4% for *12A, and 1.1% for *13) were significantly different from those reported in Caucasian populations. In the AANAT gene we found 4 novel SNPs: 2 in the 5' flanking region, 1 in exon 4, and 1 in intron 3. In the two genes belonging to the N-terminal N-acetyltransferase family, we identified 9 SNPs, 7 of them novel, for ARD1, and six novel SNPs for L1CAM. Variations at these loci may contribute to an understanding of the way in which different genotypes may affect the activities of human N-acetyltransferases, especially as regards the therapeutic efficacy of certain drugs and antibiotics.

Acetyltransferases↗

Applications of computational algorithm tools to identify functional SNPs in cytokine genes.

Understanding the functions of single nucleotide polymorphisms (SNPs) can greatly help to understand the genetics of the human phenotype variation and especially the genetic basis of human complex diseases. However, how to identify functional SNPs from a pool containing both functional and neutral SNPs is challenging. In this study, we analyzed the genetic variations that can alter the expression and function of a group of cytokine proteins using computational tools. As a result, we extracted 4552 SNPs from 45 cytokine proteins from SNPper database. Of particular interest, 828 SNPs were in the 5'UTR region, 961 SNPs were in the 3' UTR region, and 85 SNPs were non-synonymous SNPs (nsSNPs), which cause amino acid change. Evolutionary conservation analysis using the SIFT tool suggested that 8 nsSNPs may disrupt the protein function. Protein structure analysis using the PolyPhen tool suggested that 5 nsSNPs might alter protein structure. Binding motif analysis using the UTResource tool suggested that 27 SNPs in 5' or 3'UTR might change protein expression levels. Our study demonstrates the presence of naturally occurring genetic variations in the cytokine proteins that may affect their expressions and functions with possible roles in complex human disease, such as immune diseases.

Algorithms↗

Additional SNPs and linkage-disequilibrium analyses are necessary for whole-genome association studies in humans.

More than 5 million single-nucleotide polymorphisms (SNPs) with minor-allele frequency greater than 10% are expected to exist in the human genome. Some of these SNPs may be associated with risk of developing common diseases. To assess the power of currently available SNPs to detect such associations, we resequenced 50 genes in two ethnic samples and measured patterns of linkage disequilibrium between the subset of SNPs reported in dbSNP and the complete set of common SNPs. Our results suggest that using all 2.7 million SNPs currently in the database would detect nearly 80% of all common SNPs in European populations but only 50% of those common in the African American population and that efficient selection of a minimal subset of SNPs for use in association studies requires measurement of allele frequency and linkage disequilibrium relationships for all SNPs in dbSNP.

Alleles↗

TTF-1 and RET promoter SNPs: regulation of RET transcription in Hirschsprung's disease.

Single nucleotide polymorphisms (SNPs) of the coding regions of receptor tyrosine kinase gene (RET) are associated with Hirschsprung's disease (HSCR, aganglionic megacolon). These SNPs, individually or combined, may act as a low penetrance susceptibility locus and/or be in linkage disequilibrium (LD) with another susceptibility locus located in RET regulatory regions. Because two RET promoter SNPs have been found associated with HSCR, in LD with HSCR-associated RET coding region haplotypes, their implication in the transcriptional regulation of RET is of major interest. Analysis of 172 sporadic HSCR patients also revealed the presence of HSCR-associated RET promoter SNPs in LD with the main coding region RET haplotype observed in Chinese patients. By using a weighted logistic regression approach, we determined that of all SNPs tested in our study, the promoter SNPs are the most correlated to the disease. Functional analysis of the RET promoter SNPs in the context of additional 5' regulatory regions demonstrated that the HSCR-associated alleles decrease RET transcription. These SNPs overlap a TTF-1 binding site and TTF-1-activated RET transcription is also decreased by the HSCR-associated SNPs. Moreover, we identified an HSCR patient with a Gly322Ser TTF-1 mutation that compromises activation of transcription from HSCR-associated RET promoter haplotypes. Interestingly, we show that the pattern of RET and TTF-1 expression is coincident in developing human gut. We also present a detailed profile of the RET gene in our population, which provides an insight into the higher incidence of the disease in China.

Alleles↗

Optimal haplotype block-free selection of tagging SNPs for genome-wide association studies.

It is widely hoped that the study of sequence variation in the human genome will provide a means of elucidating the genetic component of complex diseases and variable drug responses. A major stumbling block to the successful design and execution of genome-wide disease association studies using single-nucleotide polymorphisms (SNPs) and linkage disequilibrium is the enormous number of SNPs in the human genome. This results in unacceptably high costs for exhaustive genotyping and presents a challenging problem of statistical inference. Here, we present a new method for optimally selecting minimum informative subsets of SNPs, also known as "tagging" SNPs, that is efficient for genome-wide selection. We contrast this method to published methods including haplotype block tagging, that is, grouping SNPs into segments of low haplotype diversity and typing a subset of the SNPs that can discriminate all common haplotypes within the blocks. Because our method does not rely on a predefined haplotype block structure and makes use of the weaker correlations that occur across neighboring blocks, it can be effectively applied across chromosomal regions with both high and low local linkage disequilibrium. We show that the number of tagging SNPs selected is substantially smaller than previously reported using block-based approaches and that selecting tagging SNPs optimally can result in a two- to threefold savings over selecting random SNPs.

Algorithms↗

Choosing SNPs using feature selection.

A major challenge for genomewide disease association studies is the high cost of genotyping large number of single nucleotide polymorphisms (SNP). The correlations between SNPs, however, make it possible to select a parsimonious set of informative SNPs, known as "tagging" SNPs, able to capture most variation in a population. Considerable research interest has recently focused on the development of methods for finding such SNPs. In this paper, we present an efficient method for finding tagging SNPs. The method does not involve computation-intensive search for SNP subsets but discards redundant SNPs using a feature selection algorithm. In contrast to most existing methods, the method presented here does not limit itself to using only correlations between SNPs in local groups. By using correlations that occur across different chromosomal regions, the method can reduce the number of globally redundant SNPs. Experimental results show that the number of tagging SNPs selected by our method is smaller than by using block-based methods.

Artificial Intelligence↗

Genotypes of SNPs of key genes regulate susceptibility and drug sensitivity to neovascular AMD in the human population.

OBJECTIVE: To compare the genetic characteristics of the normal control group to those of neovascular age-related macular degeneration (AMD) patients and to detect single-nucleotide polymorphisms (SNPs) related to the pathogenesis of neovascular AMD and the sensitivity to anti-VEGF drug, combercept. METHOD: This is a prospective case-controlled study. A total of 104 neovascular AMD patients were treated with combercept and 106 normal subjects were served as the control group. SNPs associated with neovascular AMD and disease susceptibility and drug sensitivity were analysed. RESULTS: Significant differences existed between neovascular AMD patients and normal subjects among genotypes of the SNPs of two genes, ARMS2 (rs10490924 T) and HTRA 1 (rs11200638 A). The T alleles in rs1065489 of CFH and the rs2230205 of C3 significantly promoted neovascular AMD in males while having no significant effect in females. Six SNPs of five genes, including C3 (rs2250656 G), CFB (rs2072633 G), CFH (rs2274700 A, rs3766405 T), KDR (rs6828477 A) and FZD 4 (rs10898563 T), had significant impact in reducing neovascular AMD. Two SNPs of the CFH gene (rs2274700 A and rs3766405 T) and one SNP of the CFB gene, rs2072633 G, were statistically significantly associated with good response to combercept. Conversely, the other two SNPs of the CFH gene, rs1065489 T and rs3753396 G, and the rs7412 T of the APOE gene were associated with a relatively poor patient response to drug action. Two sets of SNPs of CFB have a combined positive effect on disease. The two SNPs of CFH (rs1065489 T and rs3753396 G) and the combination of the two SNPs of CFH and rs7412T of APOE have negative effects on the drug effectiveness. CONCLUSIONS: These genotype differences facilitate the selection of individualised treatment options towards obtaining the most efficacious clinical treatment. These findings need to be validated by studies with different ethnic populations and/or larger samples.

Humans↗

A model-based approach to selection of tag SNPs.

BACKGROUND: Single Nucleotide Polymorphisms (SNPs) are the most common type of polymorphisms found in the human genome. Effective genetic association studies require the identification of sets of tag SNPs that capture as much haplotype information as possible. Tag SNP selection is analogous to the problem of data compression in information theory. According to Shannon's framework, the optimal tag set maximizes the entropy of the tag SNPs subject to constraints on the number of SNPs. This approach requires an appropriate probabilistic model. Compared to simple measures of Linkage Disequilibrium (LD), a good model of haplotype sequences can more accurately account for LD structure. It also provides a machinery for the prediction of tagged SNPs and thereby to assess the performances of tag sets through their ability to predict larger SNP sets. RESULTS: Here, we compute the description code-lengths of SNP data for an array of models and we develop tag SNP selection methods based on these models and the strategy of entropy maximization. Using data sets from the HapMap and ENCODE projects, we show that the hidden Markov model introduced by Li and Stephens outperforms the other models in several aspects: description code-length of SNP data, information content of tag sets, and prediction of tagged SNPs. This is the first use of this model in the context of tag SNP selection. CONCLUSION: Our study provides strong evidence that the tag sets selected by our best method, based on Li and Stephens model, outperform those chosen by several existing methods. The results also suggest that information content evaluated with a good model is more sensitive for assessing the quality of a tagging set than the correct prediction rate of tagged SNPs. Besides, we show that haplotype phase uncertainty has an almost negligible impact on the ability of good tag sets to predict tagged SNPs. This justifies the selection of tag SNPs on the basis of haplotype informativeness, although genotyping studies do not directly assess haplotypes. A software that implements our approach is available.

Databases, Genetic↗

Cis-regulatory variations: a study of SNPs around genes showing cis-linkage in segregating mouse populations.

BACKGROUND: Changes in gene expression are known to be responsible for phenotypic variation and susceptibility to diseases. Identification and annotation of the genomic sequence variants that cause gene expression changes is therefore likely to lead to a better understanding of the cause of disease at the molecular level. In this study we investigate the pattern of single nucleotide polymorphisms (SNPs) in genes for which the mRNA levels show cis-genetic linkage (gene expression quantitative trait loci mapping in cis, or cis-eQTLs) in segregating mouse populations. Such genes are expected to have polymorphisms near their physical location (cis-variations) that affect their mRNA levels by altering one or more of the cis-regulatory elements. This led us to characterize the SNPs in promoter (5 Kb upstream) and non-coding gene regions (introns and 5 Kb downstream) (cis-SNPs) and the effects they may have on putative transcription factor binding sites. RESULTS: We demonstrate that the cis-eQTL genes (CEGs) have a significantly higher frequency of cis-SNPs compared to non-CEGs (when both sets are taken from the non-IBD regions, i.e. regions not identical by descent). Most CEGs having cis-SNPs do not contain these SNPs in the phylogenetically conserved regions. In those CEGs that contain cis-SNPs in the phylogenetically conserved regions, enrichment of cis-SNPs occurs both within and outside of the conserved sequences. A higher fraction of CEGs are also seen to harbor cis-SNP that affect predicted transcription factor binding sites, a likely consequence of the higher cis-SNPs density in these genes. CONCLUSION: This present study provides the first genome-wide investigation of the putative cis-regulatory variations in a large set of genes whose levels of expression give rise to cis-linkage in segregating mammalian populations. Our results provide insights into the challenges that exist in identifying polymorphisms regulating gene expression using bioinformatic sequence analysis approaches. The data provided herein should benefit future investigations in this area.

Adipose Tissue↗

Replication and Functional Prediction of Two GWAS-Reported SNPs Located on RAD50 Gene Associated with Asthma in Pakistani Children.

BACKGROUND: Genome-wide association studies (GWAS) have indicated that several single nucleotide variants (SNVs) of the RAD50 gene are significantly associated with childhood-onset asthma. However, the biological role of RAD50, and its genomic variants that predispose individuals to asthma, remains unclear. This case-control study aimed to investigate the association of two Single nucleotide polymorphisms (SNPs) rs2244012, and rs6871536 of RAD50 with asthma susceptibility using experimental and computational tools. METHODS: The case-control study involved 355 participants: "176 asthma cases [mean age (sd) = 8.91 &#xb1;3.05] and 179 healthy controls [mean age (sd) = 11.10 &#xb1;8.86] from local Punjabi population of Pakistan. The SNPs were analyzed using a modified single base extension method. The allelic association with asthma and linkage disequilibrium (LD) between the two main SNPs were performed using the SHEsis tool. SNPStats was used to assess the association of SNPs under genotypic models and interaction with non-genetic factors. The LD calculator of ENSEMBL employed for the identification of proxy SNPs in high LD (r^2 > 0.97) to main SNPs. Additionally, HaploReg(v4.1) was utilized to gauge the impact of SNPs on genomic regulations. RESULTS: In current study, both SNPs were found to have a significant association (p-value <0.05) with childhood-onset asthma development under allelic and genotypic models. The alternative "G" allele of rs2244012 is shown to modify two regulatory motifs: Nrf-2 and Zbtb12, while the alternative "C" allele of rs6871536 is predicted to alter the OSF-2 motif. Moreover, 10 SNVs proximal to rs2244012 and 21 SNVs near rs6871536 are in high LD in the Punjabi population of Lahore, Pakistan (PJL). These proxy/high-LD SNVs also displayed the potential to change DNA regulatory motifs. CONCLUSION: the rs2244012, and rs6871536 variants of RAD50 gene are significantly association with childhood asthma in Pakistan. Despite being intronic variants, it is our inference that these two SNPs have the potential to either independently or synergistically regulate inflammatory responses via nearby SNVs.

Asthma↗

A tool for selecting SNPs for association studies based on observed linkage disequilibrium patterns.

The design of genetic association studies using single-nucleotide polymorphisms (SNPs) requires the selection of subsets of the variants providing high statistical power at a reasonable cost. SNPs must be selected to maximize the probability that a causative mutation is in linkage disequilibrium (LD) with at least one marker genotyped in the study. The HapMap project performed a genome-wide survey of genetic variation with about a million SNPs typed in four populations, providing a rich resource to inform the design of association studies. A number of strategies have been proposed for the selection of SNPs based on observed LD, including construction of metric LD maps and the selection of haplotype tagging SNPs. Power calculations are important at the study design stage to ensure successful results. Integrating these methods and annotations can be challenging: the algorithms required to implement these methods are complex to deploy, and all the necessary data and annotations are deposited in disparate databases. Here, we present the SNPbrowser Software, a freely available tool to assist in the LD-based selection of markers for association studies. This stand-alone application provides fast query capabilities and swift visualization of SNPs, gene annotations, power, haplotype blocks, and LD map coordinates. Wizards implement several common SNP selection workflows including the selection of optimal subsets of SNPs (e.g. tagging SNPs). Selected SNPs are screened for their conversion potential to either TaqMan SNP Genotyping Assays or the SNPlex Genotyping System, two commercially available genotyping platforms, expediting the set-up of genetic studies with an increased probability of success.

Computational Biology↗

Association studies in candidate genes: strategies to select SNPs to be tested.

OBJECTIVE: When numerous single nucleotide polymorphisms (SNPs) have been identified in a candidate gene, a relevant and still unanswered question is to determine how many and which of these SNPs should be optimally tested to detect an association with the disease. Testing them all is expensive and often unnecessary. Alleles at different SNPs may be associated in the population because of the existence of linkage disequilibrium, so that knowing the alleles carried at one SNP could provide exact or partial knowledge of alleles carried at a second SNP. We present here a method to select the most appropriate subset of SNPs in a candidate gene based on the pairwise linkage disequilibrium between the different SNPs. METHOD: The best subset is identified through power computations performed under different genetic models, assuming that one of the SNPs identified is the disease susceptibility variant. RESULTS: We applied the method on two data sets, an empirical study of the APOE gene region and a simulated study concerning one of the major genes (MG1) from the Genetic Analysis Workshop 12. For these two genes, the sets of SNPs selected were compared to the ones obtained using two other methods that need the reconstruction of multilocus haplotypes in order to identify haplotype-tag SNPs (htSNPs). We showed that with both data sets, our method performed better than the other selection methods.

Algorithms↗

Computation of haplotypes on SNPs subsets: advantage of the "global method".

BACKGROUND: Genetic association studies aim at finding correlations between a disease state and genetic variations such as SNPs or combinations of SNPs, termed haplotypes. Some haplotypes have a particular biological meaning such as the ones derived from SNPs located in the promoters, or the ones derived from non synonymous SNPs. All these haplotypes are "subhaplotypes" because they refer only to a part of the SNPs found in the gene. Until now, subhaplotypes were directly computed from the very SNPs chosen to constitute them, without taking into account the rest of the information corresponding to the other SNPs located in the gene. In the present work, we describe an alternative approach, called the "global method", which takes into account all the SNPs known in the region and compare the efficacy of the two "direct" and "global" methods. RESULTS: We used empirical haplotypes data sets from the GH1 promoter and the APOE gene, and 10 simulated datasets, and randomly introduced in them missing information (from 0% up to 20%) to compare the 2 methods. For each method, we used the PHASE haplotyping software since it was described to be the best. We showed that the use of the "global method" for subhaplotyping leads always to a better error rate than the classical direct haplotyping. The advantage provided by this alternative method increases with the percentage of missing genotyping data (diminution of the average error rate from 25% to less than 10%). We applied the global method software on the GRIV cohort for AIDS genetic associations and some associations previously identified through direct subhaplotyping were found to be erroneous. CONCLUSION: The global method for subhaplotyping can reduce, sometimes dramatically, the error rate on patient resolutions and haplotypes frequencies. One should thus use this method in order to minimise the risk of a false interpretation in genetic studies involving subhaplotypes. In practice the global method is always more efficient than the direct method, but a combination method taking into account the level of missing information in each subject appears to be even more interesting when the level of missing information becomes larger (>10%).

Apolipoproteins E↗

Automated SNP detection from a large collection of white spruce expressed sequences: contributing factors and approaches for the categorization of SNPs.

BACKGROUND: High-throughput genotyping technologies represent a highly efficient way to accelerate genetic mapping and enable association studies. As a first step toward this goal, we aimed to develop a resource of candidate Single Nucleotide Polymorphisms (SNP) in white spruce (Picea glauca [Moench] Voss), a softwood tree of major economic importance. RESULTS: A white spruce SNP resource encompassing 12,264 SNPs was constructed from a set of 6,459 contigs derived from Expressed Sequence Tags (EST) and by using the bayesian-based statistical software PolyBayes. Several parameters influencing the SNP prediction were analysed including the a priori expected polymorphism, the probability score (PSNP), and the contig depth and length. SNP detection in 3' and 5' reads from the same clones revealed a level of inconsistency between overlapping sequences as low as 1%. A subset of 245 predicted SNPs were verified through the independent resequencing of genomic DNA of a genotype also used to prepare cDNA libraries. The validation rate reached a maximum of 85% for SNPs predicted with either PSNP > or = 0.95 or > or = 0.99. A total of 9,310 SNPs were detected by using PSNP > or = 0.95 as a criterion. The SNPs were distributed among 3,590 contigs encompassing an array of broad functional categories, with an overall frequency of 1 SNP per 700 nucleotide sites. Experimental and statistical approaches were used to evaluate the proportion of paralogous SNPs, with estimates in the range of 8 to 12%. The 3,789 coding SNPs identified through coding region annotation and ORF prediction, were distributed into 39% nonsynonymous and 61% synonymous substitutions. Overall, there were 0.9 SNP per 1,000 nonsynonymous sites and 5.2 SNPs per 1,000 synonymous sites, for a genome-wide nonsynonymous to synonymous substitution rate ratio (Ka/Ks) of 0.17. CONCLUSION: We integrated the SNP data in the ForestTreeDB database along with functional annotations to provide a tool facilitating the choice of candidate genes for mapping purposes or association studies.

Algorithms↗

SNPs on human chromosomes 21 and 22 -- analysis in terms of protein features and pseudogenes.

SNPs are useful for genome-wide mapping and the study of disease genes. Previous studies have focused on SNPs in specific genes or SNPs pooled from a variety of different sources. Here, a systematic approach to the analysis of SNPs in relation to various features on a genome-wide scale, with emphasis on protein features and pseudogenes, is presented. We have performed a comprehensive analysis of 39,408 SNPs on human chromosomes 21 and 22 from the SNP consortium (TSC) database, where SNPs are obtained by random sequencing using consistent and uniform methods. Our study indicates that the occurrence of SNPs is lowest in exons and higher in repeats, introns and pseudogenes. Moreover, in comparing genes and pseudogenes, we find that the SNP density is higher in pseudogenes and the ratio of nonsynonymous to synonymous changes is also much higher. These observations may be explained by the increased rate of SNP accumulation in pseudogenes, which presumably are not under selective pressure. We have also performed secondary structure prediction on all coding regions and found that there is no preferential distribution of SNPs in a -helices, b -sheets or coils. This could imply that protein structures, in general, can tolerate a wide degree of substitutions. Tables relating to our results are available from http://genecensus.org/pseudogene.

Algorithms↗

The role of single nucleotide polymorphisms (SNPs) in understanding complex disorders and pharmacogenomics.

INTRODUCTION: In the last two years, there has been an increasing interest in single nucleotide polymorphisms (SNPs). They have been hailed as the most common polymorphism found in the human genome and are believed to be responsible for 90% of all inter-individual variation. Efforts are now directed at the large-scale identification and archiving of SNPs in the human genome. Not only are they useful markers for population divergence studies, SNPs can be utilised as markers in studies of complex diseases and pharmacogenomics. METHODS: Traditional methods for identifying SNPs, as well as methods for large-scale detection and genotyping of SNPs currently being developed, are briefly discussed in this review. Such developments will facilitate and enhance the process of identifying and characterising genes and their functions. RESULTS: The utility of SNPs in identifying genes contributing to pharmacogenetic variation and increased risk of a complex disease is discussed. The role of SNPs in influencing drug response in different individuals is also presented. CONCLUSIONS: In helping to unravel the genetic basis of complex diseases and inter-individual variation in drug response, SNPs will catalyse the transition into a new age of medicine in which medical care is tailored to the individual's genetic profile.

ATP Binding Cassette Transporter, Subfamily B, Mem↗